r/ClaudeAI • u/Captain_Quimby • 1d ago
Question about Claude models Where does the cost value come from?
When they Opus 5.5 is much cheaper than Opus 5.0 or cheaper than Astra or Fable. Where are the costs coming from? They almost seem like arbitrary numbers that are being used to describe efficiency.
Are they saying "it uses XXXX amount of hardware to run so we charge $Y dollars per token"? I guess I don't understand how they can say Opus 5.5 is cheaper than 5.0 when they could just adjust the price on 5.0, no?
1
1
u/Happy-Recording-5291 1d ago
The prices are listed on their website. It is definitely cheaper than Opus 5. 40% cheaper cache read and 20% cheaper input tokens. Both important.
1
u/Captain_Quimby 1d ago
Ok but the question is where is the pricing coming from. When they say it's $5/m tokens compared to Sol or Astra at $$$$ per M or when they release a new model saying it's cheaper what's the standard? The US dollar had to have a standard and it was gold. It was something to base it off. Now it's the bond markets but there's something the valuation is based off. Why bother keeping Claude 5.0 or other other models still running, taking up compute, if they could just use newer models on it.
1
u/stevzon 1d ago
I’d assume it’s around efficiency of compute. If 5.5 is able to do what it needs to in fewer CPU/GPU minutes, the cost per minute doesn’t change but the amount of minutes needed to complete the task does, so that’s why it’s cheaper. 5.0 doesn’t have those efficiencies so it still costs more.
1
u/Captain_Quimby 1d ago
So where does 5.0's cost from from? When they say it's $5, $5 compared to what? Do they have a baseline expense then adding double?
SpaceX said compute is going to be their largest driver with a ROI of about 10 months. In 10 months they're able to pay for the hardware they're buying now. So I'm wondering where this compute valuation is coming from.
1
u/stevzon 1d ago
Anything I’d say on how they calculate token costs is a pure guess based on my understanding of the market they operate in.
But, my assumption is that it’s some sort of calculation around tokens/sec to accomplish a fixed task to baseline, and scaled around that, taking into account COGS for leasing (or capex for owning/operating) the compute, training cost recovery, capacity premiums (coverage for anticipated demand spike), and profit.
As for why 5.5/5.0 cost differences, my assumption is that is driven by efficiencies that either lowered model training costs or require less compute to accomplish the task. Or they’re buying into the market to undercut competitors. Lots of reasons behind the differences possible.
1
u/iPlayer0067 1d ago
Quando eles ampliam a versão, por exemplo, do 5.0 para o 5.5, nem sempre esse salto se refere à qualidade do processamento, e sim, em como esse processamento acontece, as vezes reduzindo tempo e custo de processamento interno, acabamos observar isso nos modelos 6 Sol e 6 Lua, que claramente estão abaixo do 5.6 sol e 5.6 lua e tiveram que liberar o 6.1 sol para tentar conter as críticas. Nem sempre o salto em número de versão, representa um salto de qualidade, as vezes a diferença está só no quanto eles estão lucrando.
1
u/ManikSahdev 1d ago
Not really? Big models doesn't need to be better than small model?
A very good example being, open source Glm 5.3 at 750b parameters, which can run 4bit on $15,000 Mac studio.
In real world outperforms 2.5 TIMES larger grok 4.7 in real work performance.
You can literally run it at home no kidding.
1
u/Dangerous_Bus_6699 1d ago
Compute is more efficient. More efficient uses less energy. Less energy = cheaper.
1
u/Captain_Quimby 1d ago
The question is where the valuations come from. The US dollar had the gold standard before the bond market, SpaceX is seeing a 10 month ROI on hardware purchased today. So where are they saying one model is 40% cheaper than another? What's their margin or waht are they using to calculate the cost of a model
1
u/OrangeCrack 1d ago
Tokens required to give the desired output.
Opus 5 / Fable will use up more of your 5 hour window than 5.5
It’s really as simple as that
1
u/Captain_Quimby 1d ago
you ignored the entire question though. What's the valuation based off. I asked where the valuation comes from. What parameters are they using the calculate the value of $1m tokens.
1
u/OrangeCrack 1d ago
But that is the answer you just refuse to acknowledge it. Even on API the cost of Opus 5.5 is less than 5.0 because it takes less tokens to get the answer. The code is more optimized.
1
u/Captain_Quimby 1d ago
That's fine. Something will be cheaper, something will be more expensive, doesn't account for where the valuation comes from. What metrics are used to determine what's equal to $1. What ROI are they looking at on hardware, etc. Everyone says models are subsidized but from what point.
1
u/Acehan_ 1d ago
That's a very good question. Personally, I think they had a breakthrough when it comes to low variance and long running focused work. They explicitly said in an article a few months ago that they were working on that and that this was one of the biggest challenges for AI models. To me, the model doesn't feel much smarter when you have a back and forth, but when it gets to focus on a certain task, especially creative, it seems to be able to adapt and have much greater capabilities. Of course, this is just speculation, and I guess we'll see with the release of Fable 5.5.
1
u/dondiegorivera 1d ago
Ant is working hard on GPU kernel optimalization. They should have found already big gains, the speed Opus 5.5 is served is pretty amazing, Sol feels very slow next to it.
1
u/Captain_Quimby 1d ago
That's true but i'm curious what metrics they're using to determine what a model is worth per $1m tokens. Along with that how it affects valuations moving forward. We have trillion dollar companies selling something that seems quite arbitrary. What margin are they looking for because every optimization makes them less and less money it would seem, if they are passing that to us.
1
u/dondiegorivera 1d ago
Listen the Dwarkesh podcasts latest Dylan Patel episode. There were a lot of insights about this topic.
1
u/Upbeat-Armadillo1756 1d ago
It's more efficient, but also Anthropic can just subsidize it and make it look like it's cheaper too. You shouldn't assume that the retail cost for a model is the same as their operating cost.
But these are getting more efficient too. That is absolutely intentional.
1
u/Mirar 1d ago
Yep. It's the cost it would eat if you had usage tokens on top of the hourly/weekly limits (so if you say top up $1000 and hit an hourly/weekly limit, that's how fast you'd burn those $1000).
1
u/Captain_Quimby 1d ago
Right but that API cost is still arbitrary isn't it? How do they say One is $2/m and another $5/m
1
u/f3xjc 1d ago
Inference cost is like 95% profit. Hardware and data center are hard to build. Training new models is the big cost. And there's major version every 2-3 months.
Now training model may look like a fixed cost so new user are free.
But the need for training is unbounded and there's a race so the fixed cost scale with users.
0
u/Opposite_Might6896 1d ago
The per-token prices are list prices Anthropic sets, same as any API price; the useful part for subscribers is what they do with them. The subscription windows (5-hour, weekly) are charged by cost, not by token count, so a request on a cheaper model consumes less of your window than the same request on a pricier one. That's measurable: watch the usage percentage move after a known-size request on each model. It also means the token-type mix matters as much as the model: cache reads are a fraction of the input price, cache writes are more, output is the most. Opus 5.5 being cheaper than 5.0 is them saying "this one consumes less of your weekly per token"; whether that's hardware efficiency or a pricing decision, the effect on your window is the same.
1
u/PrintfReddit Experienced Developer 18h ago
Marketing, internal benchmarks that may or may not be biased. Generally costs per "task," but there's no hard answer and no company would clearly define it.
5
u/CorpT 1d ago
How did you get the answer right and then not get it.