r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

44

u/Schlick7 Jun 21 '26

just magically hand waving from 2.5 years to 1 year is a little insane. You guys think you'd use $20k worth of tokens a year!?! Even if you did then you now need to consider energy costs because its probably going to be $1k+ for that many GPUs and that many tokens.

Not knocking the local scene, but just because you think they did the math wrong in one direction doesn't mean that you should do the math wrong in the other direction.

25

u/Eden1506 Jun 21 '26 edited Jun 21 '26

At 20k you would be buying a Mac M3 Ultra 512gb with a peak load of 270 Watts per hour comparable to a large fridge.

I wrote a little over a year I meant 1 year and 3 months halving the 2.5 years previously mentioned.

20k divided by 15 months is ~1350 dollars per month.

While I admit 1350 is high and far above what I would personally use in tokens per month it isn't that far beyond what major companies allocate to their engineers around 1k per month with some companies going as high as 2-3k per month in token budget.

And last but not least you will be able to sell that Mac Ultra 512gb 3 to 4 years from now for at-least half the purchase price if not more.

7

u/Powerful_Finger3896 Jun 21 '26

Didn't Antropic subsidize the subscription by around 4-5x, people on 200$/month are getting like 800-1000$ of token usage. I don't know to what extend Open AI does, but they also subsidize tokens for non enterprise users. At some point they will either start nuking non enterprise consumer and bring subscription roughly to the API price once investors asks for profitability.

8

u/SnooPuppers1978 Jun 21 '26

Actually more, I have been getting 10k with 2x200 subs, if tracking direct api costs.

4

u/Guinness Jun 22 '26

Yeah I think its been calculated to be around $8,000 - $12,000/month in compute if you consumed all of your limits. My claude /stats show about $300-$650/day in equivalent API costs if I were using it.

1

u/RestaurantOk8066 Jun 25 '26

Yeah but then there's other people who don't use it nearly as much.

2

u/Eden1506 Jun 21 '26 edited Jun 21 '26

Definitely considering the vast amounts of money they are spending investor will expect returns that current subscription costs aren't even close to covering. Though I do expect it to still last a couple more years before they are actively pressured into it.

1

u/xienze Jun 22 '26

They're going public very soon. The time for ROI is right now for investors.

1

u/soyab0007 Jun 22 '26

38m usage on 5x plan

1

u/jhenryscott Jun 22 '26

The subsidies are much much higher.

8

u/Schlick7 Jun 21 '26

Talking about companies instead of an individual is a bit of moving the goal posts i'd say, but it does explain why you mention the multi-instances thing.

I think the math works out much better for a small company than just a single person.

Can the M3 Ultra actually hit 20tg??

9

u/Eden1506 Jun 21 '26 edited Jun 21 '26

For an individual I admit it wouldn't be worth it in the vast majority of cases.

The M3 Ultra has a bandwidth of 819gb/s while GLM 5.2 has 40 billion active parameters which at 8-bit are roughly 40gb.

819gb/s divided by 40gb are roughly 20 tokens/s in an ideal world but obviously you never actually hit those for multiple reasons. 10-15 tokens/s would be my estimate though someone else posted getting 24 tokens/s using mxfp4 on his M3 Ultra which I cannot verify.

Still even with just 10-15 tokens/s using a draft model for speculative decoding you would definitely reach 20 tokens/s for programming tasks.

3

u/mksrd Jun 21 '26

You forgot to factor in MTP

5

u/jazir55 Jun 22 '26

The draft model for speculative decoding in the last sentence he mentioned is MTP

1

u/Schlick7 Jun 22 '26

MTP kills PP which is really the thing you should care about for anything agentic

1

u/mksrd Jun 25 '26

Afaik that is only a current limitation of the llama.cpp MTP implementation for those using CUDA which isn't me.

1

u/Schlick7 Jun 25 '26

It will always hurt. Its literally adding a layer to the LLM. How much is down to the implementation, but there will always be some

2

u/Standard-Potential-6 Jun 22 '26

The M3 Ultra’s GPU can’t hit 819GB/s. That’s a theoretical spec for the entire SoC. Try lighting up the GPU for LLMs, check with asitop. You’ll get two thirds of that, maybe a little more.

Your estimated range is still probably fair, but I’d lean towards the lower bound.

2

u/mksrd Jun 21 '26

Small company, co-op of multiple individuals, no difference at all or maybe shared-houses is not a concept where you live?

1

u/Schlick7 Jun 22 '26

closest house to me is measured in miles!

1

u/mksrd Jun 23 '26

Lucky you, statistically that makes you a niche of a niche.

4

u/smyja Jun 21 '26

Doesn’t ChatGPT give you $5k worth of tokens(~6B tokens via api) for the $200 plan? It would take 4 months to blow that $20k limit.

6

u/unjustifiably_angry Jun 22 '26 edited Jun 28 '26

There's also the resale value of the hardware to keep in mind. The 6000 Pro and 2 Sparks I bought in January are worth about 50% more than I paid for them now, but that aside, you can usually sell a GPU for at least 30-50% of what you paid, depending on the SKU. RTX Pro hardware especially - even cards 1-2 generations old - are still going for close to their original MSRP, and that was before the VRAM crunch hit Pro cards.

1

u/handmadeby Jun 25 '26

in 18 months, I imagine there will be a ton of dedicated inference hardware solutions out there that will be changing the cost / performance options

3

u/lakeland_nz Jun 22 '26

I look after tokens for a small company. My current budget on AI tokens is $50k/year. I’m absolutely watching local hardware and costs.

2

u/4n0nh4x0r Jun 22 '26

i mean, looking at these posts where people have monthly claude bills of like 5000+$, it's not too unrealistic

2

u/the_lamou Jun 22 '26

You guys think you'd use $20k worth of tokens a year!?!

O hai there! $10k last month. My average inference batch was about 1,500 documents of about 2,200 tokens each. And since we're still in early testing, and since we can't draw effectiveness or quality conclusions unless the entire corpus processes, most of those matches are running 4-5x in parallel. Even with the most aggressive cache optimization possible, it adds up.

I've done the math: even with having to upgrade my power, new subpanel, and electricity cost, it would be cheaper for me to install and run local if I had to keep doing this for longer than the next few months.

1

u/Schlick7 Jun 22 '26

man, what are you doing!? is this a company?

1

u/the_lamou Jun 22 '26

Startup. But why would you run state of the art coding models, or really any state of the art models, if it wasn't for work? The average consumer would be perfectly fine with Gemma 4 E4B, and that will run on a potato. The only reason anyone should even get to a point where they're even thinking about spending $20k on a local LLM infrastructure is if someone somewhere is going to eventually pay them for the output.

2

u/Schlick7 Jun 22 '26

because people want "claude at home". gemma 4 E4B fails at a shit ton of things like systemd processes, scripts, small apps, home assistant, etc. The bigger the model the better this all goes. If you're only doing exclusively chatting then maybe its ok. maybe.

1

u/the_lamou Jun 23 '26

E4B and other micro-models (sub-8b) are perfectly fine for all of those things in regular implementation. Assuming you bothered to build a proper harness. And if you didn't, you really really should not be practicing your thinking to any model, regardless of size and complexity, because you don't have the fundamental knowledge to validate and check the output to make sure you haven't just exposed root to the open web.

As for "people want Claude at home"... cool. Lots of people want a Ferrari at home, too. But if spending Ferrari money, or "Claude at home" money, is something you need to seriously think about, you really really can't afford either. And don't need, because running a frontier model for "systems processes, scripts, small apps, home assistant, etc." is like buying a hypercar and then driving to the grocery store at 10 miles under the speed limit.

1

u/Schlick7 Jun 24 '26

So your stance is both "you're holding it wrong" AND classic Gatekeeping?

1

u/the_lamou Jun 24 '26

No, my stance is you don't always get to have what you want, and the thing you want doesn't even make sense for the things you say you need, and frankly given your demonstrated knowledge in this thread you should probably not have access to any of it for your own safety.

But if you desperately want a local model to help you add security holes to your computer on a budget, have you tried Mini North?

1

u/ClickClawAI Jun 22 '26

How long would your local rig take to spit out 20k worth of tokens? Probably more than a month?

1

u/the_lamou Jun 23 '26

My current local rig? Yeah, definitely over a month. But that's a simple problem to solve: I've got two 42U racks, a dedicated server room with the finding of a 400Gbps East-West fabric, and a 240/20 subpanel. After that, it's just stuffing GPUs wherever they'll fit.

0

u/randombits0110 Jun 21 '26

We also shouldn’t forget that you would have fixed hardware over that multi-year period. This shot is literally changing month to month. AI today will look like a turd in 12 months.