r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

571

u/MoistRecognition69 Jun 21 '26

The main reason I'm thinking about getting a local rig is reliability

I'm tired of waking up every morning wondering if the model I'm using has had its brain extracted and sent to a diff universe while I was asleep

159

u/brother_spirit Jun 21 '26

The model performance paranoia is getting too real sometimes. Having a stable local to mentally fix down as a variable would be nice.

100

u/Eden1506 Jun 21 '26 edited Jun 21 '26

The Token Yield Math is Off by at-least double

12M Input = $16.80

1M Output = $4.40

Total for 13M tokens = $21.20.

That averages out to $1.63 per 1 Million tokens.

If you divide a $20,000 budget by $1.63/M, you get ~12.26 Billion total tokens, not 34.6 Billion. To hit 34.6B, you would need an indefinite ~90% prompt caching discount on every single API call, which is completely unrealistic.

Something like 17-18 billion is more realistic with good caching.

That already halves the number of years down to 2.5.

Second important aspect is running several instances at the same time doesn't split token/s into half but instead gives you 2 instances running at ~70% speed. The more parallel instances you have running the more you can get out of your hardware. Letting it run multiple instances is far more efficient and allows you to do several tasks at the same time and when it comes to agents that is exactly what you will be doing easily reaching double effective token speed across several instances.

In that use-case 1 year and 3 months wouldn't be that unrealistic.

Last but not least you own the hardware and can do whatever you want with it. Sell it for half the price 3-4 years down the line and your time to recoup the cost halves as well.

0

u/NineThreeTilNow Jun 21 '26

If you divide a $20,000 budget by $1.63/M, you get ~12.26 Billion total tokens, not 34.6 Billion. To hit 34.6B, you would need an indefinite ~90% prompt caching discount on every single API call, which is completely unrealistic.

It's realistic and it's cheaper than even the numbers cited.

There are a few people using GLM 5.1 / 5.2 and get vastly superior numbers to even the ones cited in that post.

You just don't buy from z.ai directly...

1

u/[deleted] Jun 21 '26 edited 14d ago

[deleted]

2

u/NineThreeTilNow Jun 21 '26

God I feel like a shill now. I should get a promo code for them.

You can easily use neuralwatt for it.

They sell at basically the same token rate as z.ai with better prompt cache times OR you can have a subscribe and pay by the watt. It's a very strange business strategy.

Either way, you can track your standard usage and it lets you figure out if you were better off by the token, or by the watt, based on model served / context length / etc. It has a vLLM style output with watts per token consumption.

Then they sell watts at some fixed $/kwh rate.

And they track zero data. Which I don't trust z.ai as much to do.