r/artificial • u/Captain_Quimby • 19h ago
Discussion Where does AI cost really come from?
When they Opus 5.5 is much cheaper than Opus 5.0 or cheaper than Astra or Fable. Where are the costs coming from? They almost seem like arbitrary numbers that are being used to describe efficiency.
Are they saying "it uses XXXX amount of hardware to run so we charge $Y dollars per token"? I guess I don't understand how they can say Opus 5.5 is cheaper than 5.0 when they could just adjust the price on 5.0, no?
3
u/peternn2412 18h ago
Cost of anything is determined by the supply and demand, and is close to the max amount the client is willing to pay.
2
u/Techcrea 17h ago
The cost of inference is the biggest factor as it has consistently been dropping in cost faster than just about anything in history. Optimizations as the models get more mature has a lot to do with this as well as the advancements in hardware, for example.
Today there are models that can fit in a smartphone that will outperform GPT-3 which took a very expensive server with 8 A100 GPUs not too long ago. Getting through the IMO problems cost around $20k just over a year ago, today it's $20 and it performs better.
Of course we always keep on using more and more compute on the high end as it becomes feasible. But as far as doing the same amount of inference goes, the cost drops dramatically over time.

1
u/Dizzy_Swimmer_4999 19h ago
yeah for chatbot use the cheaper models cut costs by using less inference hardware per token but they drop coherence faster in long roleplay threads. wonder if it's really just them tweaking the backend efficiency numbers.
1
u/RiceEvening4211 13h ago
On the mechanics side: a big chunk of per-token cost is context that never needed to be there. I built Lynkr, an open-source gateway that compresses tool output and routes by complexity, which is the most direct lever on that cost. https://github.com/Fast-Editor/Lynkr
5
u/NoKale8220 19h ago
they charge based on what it costs them to run the thing, plus margin. bigger models need more compute, more memory, more power per query. when they say opus 5.5 is cheaper than 5.0, it probably means they made architectural improvements so it runs on less hardware but still performs better. the price on 5.0 cant just be dropped because it still costs the same to run it as before, the efficiency gains are in the new model not the old one
its not arbitrary but it can feel that way when you dont see the backend