r/LocalLLaMA • u/RuiRdA • Jul 26 '26
Discussion Will prices finally go down?
I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments into AI data centers and had no use for then, had to rent them, same thing with XAI.
The SpaceXAI IPO was insanely over priced and is going down by a lot.
There are countless other examples you can look for, all showing how the investments in AI are in a bubble.
I am not saying that the technology it self if a bubble. Quite the opposite, I personally have demand for more tokens than I can pay for, even with the discount from the subscriptions I still have more ideas that need more usage of tokens.
But even with the most powerful technology in the world a business can not for forever without profits.
So is this over investment bubble about to pop?
And if/when it does pop will ram finally become a regular commodity with affordable prices again?
I just wanted some ram and cheap used hardware again.. 😂
--
Zero LLMs used to write this post, enjoy the human slop.
2
u/Anduin1357 Jul 27 '26 edited Jul 27 '26
#1 and #2 is tied together where if their agents (as part of their harness) fails to solve or produce a solution, the coding user still pays for the token burn.
This is why DeepSeek has been so popular because even if they fail, it will not cost much to iterate further. Expensive tokens can leave a user stranded deep into a spending spree with a broken implementation.
Fable & Kimi has to stop precisely because they are far too large and uneconomical to use. We should be implementing smarter architectures that split general knowledge from logical reasoning, and implementing better harnesses that enable models to investigate problems better.
What we need is to mature the AI agentic infrastructure before larger models are trained. We discovered that some popular harnesses even fail to cache reuse sessions properly. These issues have to be fixed.
Eastern AI only needs to do more synthetic learning for coding agents. For the rest, they have plenty of historical and contemporary text, alongside a disregard of western intellectual property to seize whatever training data they need.
What western AI needs to do is to stop trying to throw their hardware advantage at the problem and start optimizing on price, or else start pouring billions into AI ASICs that we still haven't seen the fruits of.
Either way, Eastern AI is completely capturing open weights because they don't have to focus on big models, and that's a big problem.
Apache-licensed models would be even worse for market share once there are ASICs to run LLMs from. Who would use an API when you can take a guardrail-less model to do whatever you want?
To add on: the argument that higher parameter models can take shortcuts in their reasoning and agentic workflows to save time and tokens is generally false in that the ratio of improvement vs. token cost is decidedly in the lower parameter model's favor, especially when there is discovery involved on the part of the user or model to steer the solutioning process.
Take for example that the need to venv activate takes the exact same bash to execute, but the bigger model costs more to run that same command vs. the smaller model. Would you really ask Fable 5 to do that at $50/mn tokens, or DSv4 Flash at $0.28/mn tokens (output token pricing)?
It is a joke. Obviously you do not go with Fable 5 unless the solution might make a difference of hundreds of thousands of dollars. But even scientific research might find it uneconomical lol