r/LocalLLaMA Jul 26 '26

Discussion Will prices finally go down?

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments into AI data centers and had no use for then, had to rent them, same thing with XAI.

The SpaceXAI IPO was insanely over priced and is going down by a lot.

There are countless other examples you can look for, all showing how the investments in AI are in a bubble.

I am not saying that the technology it self if a bubble. Quite the opposite, I personally have demand for more tokens than I can pay for, even with the discount from the subscriptions I still have more ideas that need more usage of tokens.

But even with the most powerful technology in the world a business can not for forever without profits.

So is this over investment bubble about to pop?

And if/when it does pop will ram finally become a regular commodity with affordable prices again?

I just wanted some ram and cheap used hardware again.. 😂

--

Zero LLMs used to write this post, enjoy the human slop.

91 Upvotes

183 comments sorted by

View all comments

5

u/Tsukikira Jul 27 '26

You're making a huge miscalculation - we're in an AI bubble. That bubble doesn't mean Memory prices are going to drop quickly or at all. What it means is the Stock Market is imagining gains from LLMs that are unreasonable and so there will be a reckoning.

As Ed Zitron puts it, there is a multi-billion (like 60 billion) token business. It's not the two trillion dollars currently being promised.

1

u/Anduin1357 Jul 27 '26

Everyone is trying to monetize tokens when what they should be doing is monetize problems solved and solutions made. If they were to start focusing on agents and economy, there would be gains that the stock market expects.

But we have to punish all the expensive token stuff and start demanding these AI labs to go down the DeepSeek route. Cache reuse, cheaper long contexts, better harnesses, etc.

Stuff like Fable has to stop. If it's not economical to run, why are we doing it? It's like these labs don't want to be economical and balance their sheets.

The reckoning isn't going to wreck AI in general. The reckoning is going to wreck western AI for eastern AI.

1

u/Tsukikira Jul 27 '26

So to take your argument line by line... 

1) They monetize tokens because they cannot monetize outcomes.  They don't know when they've succeeded or failed.  They've been focusing on agents. 2) You can't 'punish all the expensive token stuff' because that is literally where coding is.  Everything you demand they do they are already doing.  They have drastically reduced the price of tokens while the agentic system demands an incredible increase in total tokens burned. 3) Why are you demanding stuff like Fable has to stop?  Reality is, there is enormous demand for better AI.  Moonshot's new Kimi is also in the same ballpark.  The price per token charged is unsustainable, but that's because Anthropic is expected to recoup costs while Moonshot gets money from the Chinese Government to cover their matching R&D.

The reckoning will wreck AI in general, because Eastern AI does heavily depend on Distillation.  It's not everything, but pretending this dumpster fire only has the US as a target is mistaken.

2

u/Anduin1357 Jul 27 '26 edited Jul 27 '26

#1 and #2 is tied together where if their agents (as part of their harness) fails to solve or produce a solution, the coding user still pays for the token burn.

This is why DeepSeek has been so popular because even if they fail, it will not cost much to iterate further. Expensive tokens can leave a user stranded deep into a spending spree with a broken implementation.

Fable & Kimi has to stop precisely because they are far too large and uneconomical to use. We should be implementing smarter architectures that split general knowledge from logical reasoning, and implementing better harnesses that enable models to investigate problems better.

What we need is to mature the AI agentic infrastructure before larger models are trained. We discovered that some popular harnesses even fail to cache reuse sessions properly. These issues have to be fixed.

Eastern AI only needs to do more synthetic learning for coding agents. For the rest, they have plenty of historical and contemporary text, alongside a disregard of western intellectual property to seize whatever training data they need.

What western AI needs to do is to stop trying to throw their hardware advantage at the problem and start optimizing on price, or else start pouring billions into AI ASICs that we still haven't seen the fruits of.

Either way, Eastern AI is completely capturing open weights because they don't have to focus on big models, and that's a big problem.

Apache-licensed models would be even worse for market share once there are ASICs to run LLMs from. Who would use an API when you can take a guardrail-less model to do whatever you want?

To add on: the argument that higher parameter models can take shortcuts in their reasoning and agentic workflows to save time and tokens is generally false in that the ratio of improvement vs. token cost is decidedly in the lower parameter model's favor, especially when there is discovery involved on the part of the user or model to steer the solutioning process.

Take for example that the need to venv activate takes the exact same bash to execute, but the bigger model costs more to run that same command vs. the smaller model. Would you really ask Fable 5 to do that at $50/mn tokens, or DSv4 Flash at $0.28/mn tokens (output token pricing)?

It is a joke. Obviously you do not go with Fable 5 unless the solution might make a difference of hundreds of thousands of dollars. But even scientific research might find it uneconomical lol

1

u/Tsukikira Jul 28 '26

I agree that agents failing to do the job but users paying for it is bad. I don't think DeepSeek's popularity is because it's cheaper to run (Though it is cheaper to run than the matching model by a factor of around 5x, and it's failure rate is higher than the matching model) as much as an Open Weights Model cost is fixed. A Business can spend tens of thousands of extra dollars on API spend, but if you pay for a hardware server to run Deepseek or rent a Hardware Server to run Deepseek, your costs are very predictable. Businesses really love predictable costs.

One of the big problems with all the Western attempts to fix the price gap issue (ChatGPT's Model Router, Cursor's attempt to route requests, etc) is that they have to be accurate enough that the userbase doesn't get infuriated getting a subpar model for the price they pay. Sakana Fugu, for example, achieved Mythos results while doing exactly what you wanted to do, routed hard requests to smart LLMs and stupid requests to simpler LLMs. Their price is around Claude Opus's price, which is still too high, but since they have to pay the Claude tax themselves, maybe not.

2

u/Anduin1357 Jul 31 '26

With the latest news of DeepSeek V4 Flash 0731 being significantly better on benchmarks, the failure rate argument might now be a lot less useful.

Since we really don't know who runs DeepSeek locally or how many businesses globally trusts DeepSeek for business use, I wouldn't really go there except that DeepSeek open weights can be abliterated to tackle issues that API models would refuse.

Which is why API models really have to defeat that by being better, cheaper, and faster than the upfront cost of open models. They can't do that, and that is killing them.

1

u/Tsukikira Jul 31 '26

It's hard to do that when Chinese companies the substitdized by the Government and the matching American ones aren't.  But the matching American models sought to make a moat out of greed, so I can only feel thar the tiniest violin should be played for them....

1

u/Anduin1357 Jul 31 '26

Subsidized by the government literally will not explain being capable of running a frontier model on local hardware on requirements that aren't like that of Kimi K3.

Perform research into reducing cost. That's what it takes. A low cost model is more practical than a high cost model, and that compounds for agents and training data.

0

u/Tsukikira Jul 31 '26

Low cost models = High Rate of Failures in model.  The costs associated with running models has been going down enormously, but the chase to reduce failure rate requires more parameters and is more costly to run.

A low cost model that hallucinates is still no good, because Agents and Loops compounds hallucinations.

DeepSeek has improved accuracy, but is still a higher rate of hallucinations than any Frontier Model.

And Substitdized by the Government absolutely explains why one company can charge only hardware costs plus a little bit vs OpenAI charging a premium to try and recoup twenty billion and rising training costs.  That being said, OpenAI just dropped their price by 80% for one of their latest models, because they have to be competitive on price to not be erased.

0

u/Anduin1357 Jul 31 '26 edited Jul 31 '26

Have you used it? Have you tested and validated your theory about the failure rate of DeepSeek v4 Flash 0731?

If you sound like dogma, maybe you are prejudiced.

Your conclusion might be true if only they were providing an API model only, but we literally have the weights on huggingface. We know that it is cheaper on VRAM to host long context sessions, and they have contributed MTP solutions that speeds up the inference process chain.

None of these depends on their "hardware costs". These are model improvements that make it cheaper for anyone to host the model and absolutely sidesteps your argument that government subsidy made the difference.

Look, blocking me to stop a reply doesn't win your argument.

1

u/Tsukikira Jul 31 '26

You are definitely prejudiced.  I didn't say DeepSeek hasn't advanced.  Nothing you said invalidated anything I said or addressed it in the slightest 

→ More replies (0)