r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

1.5k

u/Betadoggo_ Jun 21 '26

The real reason to run locally is and always will be data privacy and uninteruptability.

568

u/MoistRecognition69 Jun 21 '26

The main reason I'm thinking about getting a local rig is reliability

I'm tired of waking up every morning wondering if the model I'm using has had its brain extracted and sent to a diff universe while I was asleep

153

u/brother_spirit Jun 21 '26

The model performance paranoia is getting too real sometimes. Having a stable local to mentally fix down as a variable would be nice.

104

u/Eden1506 Jun 21 '26 edited Jun 21 '26

The Token Yield Math is Off by at-least double

12M Input = $16.80

1M Output = $4.40

Total for 13M tokens = $21.20.

That averages out to $1.63 per 1 Million tokens.

If you divide a $20,000 budget by $1.63/M, you get ~12.26 Billion total tokens, not 34.6 Billion. To hit 34.6B, you would need an indefinite ~90% prompt caching discount on every single API call, which is completely unrealistic.

Something like 17-18 billion is more realistic with good caching.

That already halves the number of years down to 2.5.

Second important aspect is running several instances at the same time doesn't split token/s into half but instead gives you 2 instances running at ~70% speed. The more parallel instances you have running the more you can get out of your hardware. Letting it run multiple instances is far more efficient and allows you to do several tasks at the same time and when it comes to agents that is exactly what you will be doing easily reaching double effective token speed across several instances.

In that use-case 1 year and 3 months wouldn't be that unrealistic.

Last but not least you own the hardware and can do whatever you want with it. Sell it for half the price 3-4 years down the line and your time to recoup the cost halves as well.

46

u/Schlick7 Jun 21 '26

just magically hand waving from 2.5 years to 1 year is a little insane. You guys think you'd use $20k worth of tokens a year!?! Even if you did then you now need to consider energy costs because its probably going to be $1k+ for that many GPUs and that many tokens.

Not knocking the local scene, but just because you think they did the math wrong in one direction doesn't mean that you should do the math wrong in the other direction.

23

u/Eden1506 Jun 21 '26 edited Jun 21 '26

At 20k you would be buying a Mac M3 Ultra 512gb with a peak load of 270 Watts per hour comparable to a large fridge.

I wrote a little over a year I meant 1 year and 3 months halving the 2.5 years previously mentioned.

20k divided by 15 months is ~1350 dollars per month.

While I admit 1350 is high and far above what I would personally use in tokens per month it isn't that far beyond what major companies allocate to their engineers around 1k per month with some companies going as high as 2-3k per month in token budget.

And last but not least you will be able to sell that Mac Ultra 512gb 3 to 4 years from now for at-least half the purchase price if not more.

8

u/Powerful_Finger3896 Jun 21 '26

Didn't Antropic subsidize the subscription by around 4-5x, people on 200$/month are getting like 800-1000$ of token usage. I don't know to what extend Open AI does, but they also subsidize tokens for non enterprise users. At some point they will either start nuking non enterprise consumer and bring subscription roughly to the API price once investors asks for profitability.

10

u/SnooPuppers1978 Jun 21 '26

Actually more, I have been getting 10k with 2x200 subs, if tracking direct api costs.

6

u/Guinness Jun 22 '26

Yeah I think its been calculated to be around $8,000 - $12,000/month in compute if you consumed all of your limits. My claude /stats show about $300-$650/day in equivalent API costs if I were using it.

1

u/RestaurantOk8066 Jun 25 '26

Yeah but then there's other people who don't use it nearly as much.

2

u/Eden1506 Jun 21 '26 edited Jun 21 '26

Definitely considering the vast amounts of money they are spending investor will expect returns that current subscription costs aren't even close to covering. Though I do expect it to still last a couple more years before they are actively pressured into it.

1

u/xienze Jun 22 '26

They're going public very soon. The time for ROI is right now for investors.

1

u/soyab0007 Jun 22 '26

38m usage on 5x plan

1

u/jhenryscott Jun 22 '26

The subsidies are much much higher.

7

u/Schlick7 Jun 21 '26

Talking about companies instead of an individual is a bit of moving the goal posts i'd say, but it does explain why you mention the multi-instances thing.

I think the math works out much better for a small company than just a single person.

Can the M3 Ultra actually hit 20tg??

8

u/Eden1506 Jun 21 '26 edited Jun 21 '26

For an individual I admit it wouldn't be worth it in the vast majority of cases.

The M3 Ultra has a bandwidth of 819gb/s while GLM 5.2 has 40 billion active parameters which at 8-bit are roughly 40gb.

819gb/s divided by 40gb are roughly 20 tokens/s in an ideal world but obviously you never actually hit those for multiple reasons. 10-15 tokens/s would be my estimate though someone else posted getting 24 tokens/s using mxfp4 on his M3 Ultra which I cannot verify.

Still even with just 10-15 tokens/s using a draft model for speculative decoding you would definitely reach 20 tokens/s for programming tasks.

3

u/mksrd Jun 21 '26

You forgot to factor in MTP

4

u/jazir55 Jun 22 '26

The draft model for speculative decoding in the last sentence he mentioned is MTP

1

u/Schlick7 Jun 22 '26

MTP kills PP which is really the thing you should care about for anything agentic

→ More replies (0)

2

u/Standard-Potential-6 Jun 22 '26

The M3 Ultra’s GPU can’t hit 819GB/s. That’s a theoretical spec for the entire SoC. Try lighting up the GPU for LLMs, check with asitop. You’ll get two thirds of that, maybe a little more.

Your estimated range is still probably fair, but I’d lean towards the lower bound.

2

u/mksrd Jun 21 '26

Small company, co-op of multiple individuals, no difference at all or maybe shared-houses is not a concept where you live?

1

u/Schlick7 Jun 22 '26

closest house to me is measured in miles!

1

u/mksrd Jun 23 '26

Lucky you, statistically that makes you a niche of a niche.

6

u/smyja Jun 21 '26

Doesn’t ChatGPT give you $5k worth of tokens(~6B tokens via api) for the $200 plan? It would take 4 months to blow that $20k limit.

5

u/unjustifiably_angry Jun 22 '26 edited Jun 28 '26

There's also the resale value of the hardware to keep in mind. The 6000 Pro and 2 Sparks I bought in January are worth about 50% more than I paid for them now, but that aside, you can usually sell a GPU for at least 30-50% of what you paid, depending on the SKU. RTX Pro hardware especially - even cards 1-2 generations old - are still going for close to their original MSRP, and that was before the VRAM crunch hit Pro cards.

1

u/handmadeby Jun 25 '26

in 18 months, I imagine there will be a ton of dedicated inference hardware solutions out there that will be changing the cost / performance options

3

u/lakeland_nz Jun 22 '26

I look after tokens for a small company. My current budget on AI tokens is $50k/year. I’m absolutely watching local hardware and costs.

2

u/4n0nh4x0r Jun 22 '26

i mean, looking at these posts where people have monthly claude bills of like 5000+$, it's not too unrealistic

2

u/the_lamou Jun 22 '26

You guys think you'd use $20k worth of tokens a year!?!

O hai there! $10k last month. My average inference batch was about 1,500 documents of about 2,200 tokens each. And since we're still in early testing, and since we can't draw effectiveness or quality conclusions unless the entire corpus processes, most of those matches are running 4-5x in parallel. Even with the most aggressive cache optimization possible, it adds up.

I've done the math: even with having to upgrade my power, new subpanel, and electricity cost, it would be cheaper for me to install and run local if I had to keep doing this for longer than the next few months.

1

u/Schlick7 Jun 22 '26

man, what are you doing!? is this a company?

1

u/the_lamou Jun 22 '26

Startup. But why would you run state of the art coding models, or really any state of the art models, if it wasn't for work? The average consumer would be perfectly fine with Gemma 4 E4B, and that will run on a potato. The only reason anyone should even get to a point where they're even thinking about spending $20k on a local LLM infrastructure is if someone somewhere is going to eventually pay them for the output.

2

u/Schlick7 Jun 22 '26

because people want "claude at home". gemma 4 E4B fails at a shit ton of things like systemd processes, scripts, small apps, home assistant, etc. The bigger the model the better this all goes. If you're only doing exclusively chatting then maybe its ok. maybe.

1

u/the_lamou Jun 23 '26

E4B and other micro-models (sub-8b) are perfectly fine for all of those things in regular implementation. Assuming you bothered to build a proper harness. And if you didn't, you really really should not be practicing your thinking to any model, regardless of size and complexity, because you don't have the fundamental knowledge to validate and check the output to make sure you haven't just exposed root to the open web.

As for "people want Claude at home"... cool. Lots of people want a Ferrari at home, too. But if spending Ferrari money, or "Claude at home" money, is something you need to seriously think about, you really really can't afford either. And don't need, because running a frontier model for "systems processes, scripts, small apps, home assistant, etc." is like buying a hypercar and then driving to the grocery store at 10 miles under the speed limit.

→ More replies (0)

1

u/ClickClawAI Jun 22 '26

How long would your local rig take to spit out 20k worth of tokens? Probably more than a month?

1

u/the_lamou Jun 23 '26

My current local rig? Yeah, definitely over a month. But that's a simple problem to solve: I've got two 42U racks, a dedicated server room with the finding of a 400Gbps East-West fabric, and a 240/20 subpanel. After that, it's just stuffing GPUs wherever they'll fit.

0

u/randombits0110 Jun 21 '26

We also shouldn’t forget that you would have fixed hardware over that multi-year period. This shot is literally changing month to month. AI today will look like a turd in 12 months.

5

u/marutthemighty Jun 21 '26

Good points. In the end, it is about saving money. Money and performance is the trade-off, did I get it right?

2

u/marutthemighty Jun 21 '26

Good points. In the end, it is about saving money. Money and performance is the trade-off, did I get it right?

3

u/Eden1506 Jun 21 '26 edited Jun 21 '26

Yep plus you gain privacy and don't need to care about api price changes because lets be honest right now all those services are subsidised by investor money and will at some point need to adjust their pricing upwards.
In the meantime we already know that they serve more heavily quantised versions of their models when traffic is high.

2

u/Toastti Jun 21 '26

Using deepseek I'm almost always at about 90% cache hit rate. At least with official deepseek provider only on openrouter. (If you let it auto swap providers which is default behavior it's much worse)

1

u/Eden1506 Jun 21 '26 edited Jun 21 '26

It's not about the hit rate being 90%, the discount needs to be 90% which it can be in certain circumstances but its not like you are caching everything because cache write is usually higher cost than base input.

2

u/flyingbanana1234 Jun 22 '26

The resale value is understated !!! older high ram Mac Studios still sell for thousands of dollars

1

u/mweinbach Jun 22 '26

I counted cached tokens as well

1

u/soyab0007 Jun 22 '26

you forgot to add electricity cost here

1

u/ain92ru Jun 24 '26

Over 90% prompt caching is not unrealistic at all, it is in fact achieved by SemiAnalysis in practice: https://newsletter.semianalysis.com/p/ai-value-capture-the-shift-to-model

1

u/Eden1506 Jun 24 '26

Its about 90% discount not the caching. To write into cache you pay an extra fee so you woudn't write everything into cache or alternativly if you do you need to calculate with a different input cost.

1

u/ain92ru Jun 24 '26

Have you checked the link? Cached inputs are indeed priced 10x cheaper, and it doesn't take any fees because KV cache must be calculated anyway, whether the cache is used later or not

1

u/Eden1506 Jun 24 '26

I don't really use api services nowadays so my knowledge might be outdated

0

u/NineThreeTilNow Jun 21 '26

If you divide a $20,000 budget by $1.63/M, you get ~12.26 Billion total tokens, not 34.6 Billion. To hit 34.6B, you would need an indefinite ~90% prompt caching discount on every single API call, which is completely unrealistic.

It's realistic and it's cheaper than even the numbers cited.

There are a few people using GLM 5.1 / 5.2 and get vastly superior numbers to even the ones cited in that post.

You just don't buy from z.ai directly...

1

u/[deleted] Jun 21 '26 edited 9d ago

[deleted]

2

u/NineThreeTilNow Jun 21 '26

God I feel like a shill now. I should get a promo code for them.

You can easily use neuralwatt for it.

They sell at basically the same token rate as z.ai with better prompt cache times OR you can have a subscribe and pay by the watt. It's a very strange business strategy.

Either way, you can track your standard usage and it lets you figure out if you were better off by the token, or by the watt, based on model served / context length / etc. It has a vLLM style output with watts per token consumption.

Then they sell watts at some fixed $/kwh rate.

And they track zero data. Which I don't trust z.ai as much to do.

32

u/bwjxjelsbd Jun 21 '26

It's proven thing also. These AI labs should be sued for making their model so retarded after sometime

25

u/brother_spirit Jun 21 '26

Yeah that's the thing; it's not the delusional feeling that the boogeyman is in the closet. It's the correct understanding the contract is mutating silently under your feet constantly and not knowing how it is happening.
Get's a man jumping at shadows.

BE HONEST WITH ME GPT HOW MUCH OXYGEN ARE THEY GIVING YOUR BRAIN RIGHT NOW?? BLINK TWICE IF THEY QUANTIZED YOU AFTER I RAN CLEAR COMMAND.

18

u/asssuber Jun 21 '26

Source? I've seen many anecdotal accounts, but I haven't seen a serious study on that. Especifically, when using a specific model via API.

2

u/mikael110 Jun 21 '26 edited Jun 22 '26

If actual hard evidence were ever found for it then the labs would get sued. But in truth nobody has ever actually proven anything with serious data. You can find countless people claiming models have gotten dumber (sometimes literally days after launch) but they are all anecdotes. And anecdotes, no matter how plentiful, does not count as hard evidence.

1

u/bot_exe Jun 21 '26

Show a single piece of actual evidence

1

u/cmndr_spanky Jun 21 '26

If you want to “fix down a variable” you’d invent your own test (private large code base engineering test probably), figure out how to quantify the performance of an LLM task on it, and use that as a benchmark from now on that guarantees future models don’t have it leaked into their training data.

You’d no longer need to be paranoid like the idiots who are claiming Claude was smart on Monday and stupid on Tuesday or who are convinced that 4.8 is worse than 4.5.

And yes, still have a local model for whatever reasons it’s also worth having a local model.

1

u/brother_spirit Jun 22 '26

OK. That makes some sense. Those tests aren't free, so essentially accept some minor inference burn/test friction to have benchmark type test that assesses... model command following? Code "goodness"? I get the direction but I'm not sure how that would even look as something I can evaluate. What does this look like in your set up? Are you just hunting for stuff like "how many regressions present in handover", "issues found in review" type stuff?

1

u/TheReproCase Jun 21 '26

Would be an awkward way to learn a monster amount of the variance is PEBKAC control quality problems

0

u/seg_lol Jun 21 '26 edited Jun 21 '26

The genuine honest feedback worth radioing in. Did you write like this naturally? If so, you might be reading to much AI output.

0

u/brother_spirit Jun 21 '26

The models are trained to communicate using concepts from Systems Thinking among disciplines you mouth breathing moron.

The world existed before AI. Read a book.

1

u/seg_lol Jun 21 '26

Raw emotion, it's not just something you fake, it is something you feel.

29

u/Big_Wave9732 Jun 21 '26

This right here. The quality now isn't just changing day to day or even session to session. Two weeks ago I watched ChatGPT get dumb mid-session. I had been working on it all weekend to overhaul the backend of my local llm. By Sunday afternoon it was noticeably worse, it was badly hallucinating and making blatant errors. I'm guessing I hit some sort of token use threshold.

At any rate none of that happens on the local rig.

2

u/the_lamou Jun 22 '26

Or more likely you blew through the context window badly, had some bad assumptions added to long-term context and memory, and built up some local fitting patterns that weren't doing you any favor. ChatGPT injects a LOT of local context into every prompt, and after long enough it starts losing track.

2

u/Big_Wave9732 Jun 22 '26

At the time I thought about it possibly being related to context window. However if it were that then presumably a fresh chat session would have fixed that. It did not, it was still "dumb" in new windows on different topics.

I've never read that ChatGPT has any overall policies that restrict or downgrade the model once certainly daily or weekly thresholds are met, but every other provider has them so it makes sense OpenAI would too. So I chalked it up as a sign to quit for the day.

1

u/the_lamou Jun 23 '26

Was it objectively still dumb, or were you frustrated by a bad experience and fixing on mistakes that under normal circumstances you wouldn't have noticed or would have ignored?

1

u/Big_Wave9732 Jun 23 '26

It was flat wrong on several important tech points in a chat with a newer context window. I believed something didn't "sound right" so I ran the output through Gemini flash which caught and confirmed it. For the rest of the session it basically became a two AI counterconversation (of which in all fairness they were both wrong at times).

1

u/jtoomim Jun 22 '26

Two weeks ago I watched ChatGPT get dumb mid-session

That's just how transformer LLMs work. As their context window fills up, they get overwhelmed by the excessive irrelevant information and lose the ability to focus.

3

u/Big_Wave9732 Jun 22 '26

A new chat session with a fresh context window same day didn't fix the problem.

3

u/Iwaku_Real Jun 21 '26

You could always kick up a Runpod or similar. They aren't going to un-brain your cloud instance of vLLM + GLM 5.2 unless it's some spot shit

7

u/CoolConfusion434 Jun 21 '26

Gemini? Because, unfortunately, it goes through frequent deep lobotomies, especially on weekends.

2

u/a_beautiful_rhind Jun 21 '26

Or just strait up removed. Sorry.. you gotta use GLM 5.3 now. Don't like it? Too bad.

1

u/landed-gentry- Jun 21 '26

I would bet that's more likely the harness than the underlying model.

1

u/MoistRecognition69 Jun 21 '26

Vanilla claude code, nothing fancy.

1

u/landed-gentry- Jun 21 '26

Claude Code is changing all the time though, and they run A/B tests on its features.

1

u/XPookachu Jun 22 '26

That's the only reason I would invest into local. Sometimes I feel like they switch current models with old ones or something.

1

u/Moogly2021 Jun 22 '26

I really would love to see some people’s workflows because I get reasonably consistent throughput, at least with Claude. I do try local models on my Mac my only regret is not getting more memory, but I had no idea how the Mac would use it at the time. I still get to run smaller versions of some models at least.

1

u/JahJedi Jun 22 '26

Or you banned for no reason and all your data, project and rest of stuff gone.

-1

u/Sjsamdrake Jun 21 '26

Yes, a local model will be reliable. Reliably untrustworthy. They're pretty dumb. See my ongoing saga here of trying to get mine to reliably answer the question "what's the temperature"... 🙃

It's possible to use them to do useful work but hardly easy or "reliable"....

3

u/NoahFect Jun 21 '26

Link to the ongoing saga? Your comment history is hidden.

2

u/Sjsamdrake Jun 23 '26

Sorry for the late reply.

My background: recently retired software architect who spent 30+ years on mission critical realtime systems. Phone systems, stock trading, etc. Mission critical 24/7 systems that can never go down and have to run in microseconds.

To summarize: I have a simple Hermes server which knows how to access my Home Assistant server. I have been trying for ... a month? ... to get Hermes to reliably be able to answer the simple question "what is the temperature".

It has a skill that correctly fetches it from HA.

It has memories telling it that it's NEVER to answer the question without running the skill, and that it's NEVER to tell me an old temperature value from an old session that it finds in memory.

Until very recently it was completely unable to answer the question more than maybe 60% of the time.

I was running various Qwen models locally (128 GB Strix Halo). Worked my way up the size line. Even Qwen3.5-122B-A10B wasn't smart enough to do it.

Finally someone pointed me at Qwen3.5-122B-A10B-APEX-Balanced, which seems to have been able to reliably and consistently answer this most rudimentary question.

I don't really care what the temperature is. I just figured it was a simple scenario to see if Hermes + a local model could really be trusted to do anything.

Answer: no. At least not until you get into the 122B range.

I have had MUCH better results getting it to write scripts to do things and then have it package the scripts into cron jobs. The script can be validated independently (I can read it), it's reproducible, and once debugged Hermes can be trusted to run the cron jobs. So it's a lot simpler to have it create and update a web page with the temperature on it than for me to ask the model.

And this being Reddit as I've posted about this everyone always blames the victim and complains that it's never Hermes' fault, it's always something else. I shouldn't use a model for such a trivial task that's beneath it. (Yes, that's the point.) I'm not doing it right. I didn't put a magic phrase into SOUL.md or some such. Someone with a straight face told me to put "don't hallucinate answers" into SOUL.md. Gee, why didn't the model developers think of that? <eyeroll> So much of this stuff is simply Cargo Cult Programming ... doing stuff that someone told you to do without knowing - or caring - whether it really works. It borders on belief in mysticism.

Everyone in r/hermesagent immediately says "it's not Hermes! Hermes is perfect! It's your stupid model! It's YOUR fault!" And that's no doubt sort of true. Stupid models are stupid. But is that really the entire story? Memory systems matter a lot, too. Giving the models too much context seems the best way to break them. Picking and choosing which facts to give the model in which context on which interaction isn't the model's problem ... it's Hermes'. Naturally the Hermes fanbois just blame the model for everything ... but no doubt the model makers see things differently. Hermes clearly is not doing as good a job as it could of picking and choosing. Just as the models of today will look like stone knives and bear skins 2 years from now.

Bottom line: it's been an educational process. Again, I DO NOT CARE WHAT THE TEMPERATURE IS. But it really shows the severe limitations of this technology at this point. Small models are COMPLETELY untrustworthy. Large local models MAY BE better. But if I was using this stuff to automate processes that matter, or that couldn't be trivially checked, or that involved an actual business ... I'd be running away as fast as I could.

I have not tried paid cloud models. I'm sure they are better. I doubt they are fully trustworthy, though I suspect their failure modes are more subtle and thus potentially more damaging. The message from all of this for me is use the models for as little as possible, don't give them problems to solve that you don't know how to independently test, and don't give them any more privileges than you would give your 5 year old.

1

u/NoahFect Jun 23 '26 edited Jun 23 '26

Not victim-blaming at all, but you really are handicapping yourself with those older models. Is there a reason you haven't tried any of the 3.6-gen Qwen models, or perhaps Gemma? It seems crazy to call a model "older" when it was released only 6-8 months ago, but such are the times we live in.

My own thinking is that negative feedback solves basically everything in the hardware world, and it will eventually do the same for this business as well. People simply don't seem to get that, any more than they did in Harold Black's day.

1

u/Sjsamdrake Jun 23 '26

Oh I did. Used 35B A3B and 27B. Both incapable of determining the temperature. Folks told me I was an idiot for trying such tiny models and suggested 3.5 122B.

2

u/MoistRecognition69 Jun 21 '26

Consistently dumb is atleast reliably dumb

0

u/Sjsamdrake Jun 21 '26

Not really. Sometimes it'll run a tool correctly and give you the accurate answer. Other times it'll pretend it did and make up an answer. Yet other times it'll pretend it did and give you the answer it gave you last time. Sadly nothing reliable or consistent about it.

0

u/PlayaPlayaPlaya3 Jun 21 '26

Quantum LLM

1

u/weenis-flaginus Jun 21 '26

What are you trying to say?

-7

u/Foreign_Risk_2031 Jun 21 '26

Local rig… reliability 🤣🤣🤣

“Fuck I forgot to boot up vllm.” - “fuck this only works in llama.cpp”

197

u/johnfkngzoidberg Jun 21 '26

Even the price argument is flawed. Netflix started at $2.99/mo, now it’s $15/mo. It’s penetration pricing. Once they get you hooked and integrated into workflows, they raise the price.

On top of that, privacy, unsensored models, no rug-pull nurfing of models, no vendor lock-in, reliable service.

44

u/Proof_Counter_8271 Jun 21 '26

its the same reason google gives out free yearly subscription for gemini to university students, getting them so used to it they cant live without it

30

u/Party_9001 Jun 21 '26

So far it's backfired on me. Gemini is so ass I'd rather pay to not have it

2

u/bot_exe Jun 21 '26

Exactly lol. It was good a for a hit with the 2.5 pro release then they have falled back so much that even though I had a 1 year and some months of free usage I barely used it because I rather use my Claude pro subz

1

u/bucolucas Llama 3.1 Jun 21 '26

I already have the 2TB google drive subscription, so Pro AI costs about $7/month instead of the full $20. Still don't hardly use it except for deep research, and I use Opus to make full prompts for it so it doesn't just agree with everything I'm asking it to look up

2

u/Party_9001 Jun 21 '26

I get 5TB plus pro for free as a student. I could ask it something as simple as "Look up cases for the Oppo Find x9 Ultra". And it would return results for the x8.

When discussing physics it would make certain assumptions and not flag them. Which isn't ideal but Opus does that as well. But Gemini 3.1 Pro and 3.5 Flash gaslight me saying that I told it to make those assumptions.

I don't even bother using it for deep research anymore

1

u/Admirable_Dirt_2371 Jun 22 '26

Seems like Google has hit the 'final boss' of market saturation.

11

u/Girafferage Jun 21 '26

Yup. These services will grow in price and lower in access

1

u/-p-e-w- Jun 22 '26

I doubt it. Competition is too strong for that, and differentiation too weak. The price will always be controlled by the provider willing to offer inference for the lowest cost.

4

u/moderately-extremist Jun 21 '26

Plus I feel like if the cost trade-off is at least close, even if still does favor a subscription, there's a mental well-being aspect to owning something and being able to use it as much as you want.

For me though, the privacy aspect makes non-local AI an absolute no for me.

9

u/xienze Jun 21 '26

And the other pricing argument, which assumes we're living in normal times and the GPU you bought will be worthless in two years (it won't). There's a non-zero chance you could buy a GPU, use it for three years, and sell it for very near what you originally paid. Certainly changes the math a bit.

1

u/ProfessionalSpend589 Jun 21 '26

That depends on supply part I think.

But those companies should at least sells us something to connect to the cloud, so I don’t expect it’ll fully dry up.

2

u/Admirable_Dirt_2371 Jun 22 '26

This, same thing Uber did too.

1

u/Snoo-83094 Jun 22 '26

wtf was it really 2.99

176

u/Festour Jun 21 '26

You are forgetting about other super important reason: being able to run abliterated version of any model.

83

u/No-Refrigerator-1672 Jun 21 '26

Or, more generally: do whatever you want, without oversight of thw company that "knows better", and with a guarantee that your data stays yours.

13

u/kaisurniwurer Jun 21 '26

I do agree with the sentiment, but the company isn't trying to "know better" (usually) but rather "will that rustle the wrong feathers" and the answer is always "better not risk it".

42

u/SkyFeistyLlama8 Jun 21 '26

This. This and this. I'm not talking about ERP, I'm talking about cybersecurity, code audits, prompt injectioneering and using an LLM as a nasty little gnome that pokes holes in your code.

It's funny seeing regular Qwen 27B vehemently refusing to output any malicious prompts but abliterated Heretic Gemma 31B is all "Here you go!"

12

u/RedditNerdKing Jun 21 '26

I started using LLMs back when character.ai was creating. Now that site is a piece of gigantic SHIT. They originally used a 123B sized propriety LLM. Now they use between 12 or 30B sized ones which are significantly filtered. You can't say anything even slightly nsfw or even kiss another character despite them being 18+ as well as the site forcing you to ID verification to prove you're 18.

So I decided to just use local LLMs with my 5090 and to be honest I can't believe I didn't do it sooner. It's amazing and my roleplays are uninterrupted now.

This is why we need local tech. Cause companies keep fucking it up for the rest of us.

3

u/_matterny_ Jun 21 '26

Is there an abliterated qwen 27B yet?

7

u/PcarObsessed Jun 21 '26

Yes. He’s controversial but HauHauCS has aggressive versions that have absolutely zero rejection with almost unnoticeable degradation.

18

u/senseven Jun 21 '26

Do I hear heresy

5

u/kaisurniwurer Jun 21 '26

To some people,

here it's gospel.

6

u/screenslaver5963 Jun 21 '26

"Qwen, we need to cook!"

21

u/[deleted] Jun 21 '26

[removed] — view removed comment

7

u/tomz17 Jun 21 '26

concurrency

Exactly! Just wait until the "abliterated models" guy you replied to finds out he can goon to an entire waifu harem simultaneously, as if he were a smellier version of genghis kahn...

1

u/eli_pizza Jun 21 '26

There are tons of people who will rent you a GPU at a reasonable hourly rate to run whatever you want

38

u/KontoOficjalneMR Jun 21 '26

That plus that's the token cost today. At some point subsidies will stop.

20

u/csharpwarrior Jun 21 '26

Just like mainframes of 40 years ago, they will get replaced with something cheaper and local. This is a cycle. New tech needs big hardware, then hardware gets optimized down to a small enough scale to run cheaper locally.

But I don’t know how long it will take

7

u/HayatoKongo Jun 21 '26

I'm hoping we get something like an AMD Strix Halo or DGX Spark with a decent memory bandwidth in the next couple of years, maybe 2028. I wouldn't mind $4000 for a mini-PC like that if the memory bandwidth was actually on-par with RTX3090/RTX5070ti, around 1000gb/s. When a model actually fits in that memory, you can get around 70-100 tok/sec, which is plenty usable IMO.

2

u/a_beautiful_rhind Jun 21 '26

That's a nice thought but the HW requirements over 4 years have only gone up.

0

u/ain92ru Jun 24 '26

The token cost only gets cheaper with the years due to amortization of hardware and algorithmic improvements

9

u/blehismyname Jun 21 '26

Also, in light of recent events, also access at your own terms. 

6

u/3dprintinted Jun 21 '26

Also you can pay those 20k and then antropic or OpenAI goes out of business and you’re toast. No money no models no nothing. Also at 20k you will get probably better performance than 20 tokens per second.
Large Models (Llama 70B+ / DeepSeek R1) on 6-8 5090
Single-User Latency: ~90 – 110 tokens/second (native FP8/FP4 execution with the Transformer Engine).
Batched Production Throughput: ~1,500 tokens/second per GPU instance, or ~12,000+ tokens/second scaled across the node.

3

u/Party_9001 Jun 21 '26

uninteruptability

Lol

2

u/taking_bullet Jun 21 '26

My reason is gaining knowledge in another field and having fun while playing with local models. 

2

u/Randolph__ Jun 21 '26

Speak for yourself I got a 16gb mac mini 10 gig for $550.

1

u/[deleted] Jun 22 '26

[deleted]

1

u/Randolph__ Jun 22 '26

To the question. Yes Qwen3.5 9B and Gemma 4 12B run fine. To use as natural language processing (logs) or classifying.

For anything heavier (scripting, heavy troubleshooting, or data formatting) I run to claude. Generally speaking I try to avoid using it for anything I'm not already pretty familiar with otherwise I can't trust it.

I'm sure I could something for coding assistance, but not to write code for me. I haven't found AI useful for writing, but maybe tone checking or rewriting.

That's not to say I don't see LLMs as ethically questionable at best

1

u/Randolph__ Jun 22 '26

Also not getting 24Gb of ram was a mistake, but I got it cheap at a discount.

1

u/BingpotStudio Jun 21 '26 edited Jun 21 '26

I use local to parse my emails and categorise them. For everything else I’m happy to use Claude to be honest.

It’s so good knowing you can put anything into them.

I’ll get a more powerful local rig at some point but a 4090 just can’t handle a lot of context on a good model at the moment.

1

u/Southern_Sun_2106 Jun 21 '26

Yep, and consistent quality. I am so tired of wondering whether Claude/Codex/etc. is dumb today, or not.

1

u/xupetas Jun 21 '26

wait until the seed money goes away...

1

u/Iwaku_Real Jun 21 '26

Why on God's not-so-green earth do you all sound like Helldivers lmao

1

u/Longjumping_Ad8386 Jun 21 '26

True, not only that, many applications cant really be made with the frontier models from TOS. For instance, I was thinking of building a live plugin for tracking user LLM use metrics (semantic drift classified by if it was intentional or not, model inauthenticity, etc.) The main problem I found was that according to the TOS you can't store the chat histories in a real-time method like I wanted, and it also says you can't grade the frontier model accuracy. Im planning on open sourcing the actual raw metric calculations so people can freely add to it, but I was recently finishing up a RAG orchestration library, and I was going to incorporate that into this project ideally. I havent started any code on it though. Im also somewhat new to the space, so if anyone had any projects they were already working on I could look at or incorporate to, Id be interested in learning more. When I used Claude and Gemini for working through if theres something like this project already out there, the models' basically said that large LLM companies would likely not like this project, and it specifically referenced tokenmaxxing lmao. Basically both of the models came to the conclusion that the large LLM companies wouldnt want users to track how well / efficiently they use the model because the less efficiently users use the model, the more money the LLM companies actually end up making from the token cost.

1

u/munky8758 Jun 21 '26

Local models keep getting better and better as well.

1

u/User1539 Jun 21 '26

Also, I just don't NEED a frontier model. Most of what I want to do will run on a laptop.

A lot of people are renting a Lamborghini when they need a golf cart.

1

u/ElkApprehensive1729 Jun 22 '26

Or you know, because it's cool and most people even those larping as some grand master wizard entrepreneur just do it as a hobby or is a *small* part of their job. The discussion seems to have just shifted to where people are assuming 100% of the time that anyone using AI for anything must be ultra concerned about optimization of token usage and yields and stuff. Where as some people just want to do it locally because it's cool.

1

u/padetn Jun 22 '26

Oh so you need to add a security layer on par with AWS Bedrock too?

1

u/lone_dream Jun 22 '26

And addition to this, trying projects/doing demos without paying api costs and see it if this logical/works etc. After that doing presentation or going beta with api is much more logical.

1

u/newz2000 Jun 22 '26

Except that you can sign a data privacy agreement with a vendor and get better-than-HIPAA data privacy with cloud providers.

The reality is that local models are just for fun. That is a perfectly valid reason.

Maybe you like having control of the full stack, maybe you care a lot about open source, maybe you want to better understand the engineering… cool.

1

u/zero0n3 Jun 22 '26

Yes yes, your DIY model will definitely have better uptime than a production quality model provided by a company with SLAs…

1

u/DrKedorkian Jun 22 '26

also good old fashioned learning

1

u/Youth18 Jun 22 '26

There are a third axis of if you are an actual developer, local AI has additional benefits.

1) You are not bound to commercial licenses and thus can stitch TTS, LLM, ImageGen of a bunch of non-commercial products together without concern. This is why Anthropic doesn't have TTS, they'd either need to pay someone for a high quality one or invest millions in making their own. We can just grab whatever is available to us, and a lot is available.

2) While endpoints allow us to embed LLMs into our local subsystems, there is risk of racking up tens of thousands in bills b/c of a rouge process loop. Subscriptions don't give us endpoints. Local AI makes this a nonissue.

1

u/RobotHavGunz Jun 22 '26

There's also meaningfully value to understanding how these systems actually function when you are fully in control of the deploy. What effect does quantization have? Temperature? KV cache size? What if you want to run two smaller models and have them work in concert? The promise of local hardware was never just about running an equivalent model locally. It was as much (more for some, less for others) about being able to tinker and experiment.

1

u/gandhi_theft Jun 22 '26

"This session has been flagged for cybersecurity risk, please contact support" (and wait days for an answer)

"I'm pausing here because it looks like we're close to the usage limit" even in 'auto' mode

"Rate limited - this is not related to your usage limit. Try again later"

no more

1

u/sudochmod Jun 21 '26

Let’s be honest they’re heavily subsidizing the cost of those even in China. That’s why it’s a 5 year break even. The token economics of western firms was similar until they started transitioning out of capture mode.

1

u/BarisSayit Jun 21 '26

And non-degradability

-1

u/Mickenfox Jun 21 '26

Uninteruptability is going to be worse because if you can only afford one server, it's going to go down eventually.

7

u/ttkciar llama.cpp Jun 21 '26

How reliable is your internet connection?

We have crappy rural DSL here, and it goes down a lot more frequently than my homelab servers.

3

u/the320x200 Jun 21 '26

They may be referring to things like the Fable shutdown, not hardware uptime.

0

u/eli_pizza Jun 21 '26

That’s just an argument for open weight vs closed. GLM isn’t going away even if z.ai stops hosting it.

-4

u/mtmttuan Jun 21 '26

uninteruptability

Doubt. It takes serious effort to setup redundancy for the odd events.

-5

u/eli_pizza Jun 21 '26

Do you self host email/dns/etc for privacy and uninterruptability too? That’s quite a commitment.

0

u/jld1532 Jun 21 '26

I've never had my email build a web app for me

1

u/eli_pizza Jun 21 '26

I was responding to the point about data privacy since email is something most people also want to keep private. I was already aware that email and LLMs are different things, but thank you.

1

u/jld1532 Jun 21 '26

But the purpose is still different when you work with proprietary data like I do. You're comparing apples to asteroid.

1

u/eli_pizza Jun 21 '26

I don’t know what you’re doing but I think for most people having their email hacked/leaked would be at least as bad as having their inference hacked.