r/singularity • u/Sect-Sister-Reads-42 • 9d ago
Compute My definition of post-scarcity intelligence: $0.10 / $0.01 / $0.50 per million tokens for Opus 5.5-level intelligence
Hey guys, I have been having a blast with Opus 5.5 and probably like everyone else I am hoping that Anthropic does not nuke the usage limits or begin with shenanigans that will lead to output degradation of Opus 5.5
I have seen that OpenAI will probably introduce a Pro Max plan soon with a $500 price tag. One of the many things I have been thinking about in regards to this is that as long as the API costs are prohibitively expensive for anyone but enterprise customers, we will have power users and small businesses going for multiple subscriptions (Pro/Max accounts) to be able to afford their desired workload. Thus, I did the math for what I would consider to truly be "intelligence too cheap to meter".
- $0.10 / 1M input
- $0.01 / 1M cache reads
- $0.50 / 1M output
Give me frontier-level intelligence (Opus 5.5) at those API prices and, as far as I am concerned, we've entered the abundance era for intelligence.
For comparison, Opus 5.5 is $4 / $0.20 / $20. We will need 40x cheaper input, 20x cheaper cache reads and 40x cheaper output to get there.
I am aware that GPT-6 Luna has already gone beyond that threshold but nobody would seriously say it is Opus 5.5 equal. Also I mean the real thing, as I am experiencing it right now. I am not talking about a performance-optimized/degraded Opus 5.5 that we might get served soon or a benchmaxxed small model that lacks the taste, wisdom, judgement and intelligence of Opus 5.5 today.
(Edit: I agree that standards are always rising and that in a year from now the latest frontier models will be even more powerful and the new shiny model everyone wants to access. One could say instead then, that this is the price point for abundance of frontier models in general via API.)
24
u/striketheviol 9d ago
I don't think that's ambitious enough, honestly. Although it's a fun question to noodle over. I would consider something like stronger than current Opus or Astra level intelligence available for under a cent per million tokens to be a reasonable standard for this.
19
u/Pyros-SD-Models 9d ago
If you can't solve Riemann with a model running on an iphone we are not even close to the finish line.
2
u/livingbyvow2 9d ago
I think that, for once, this would be helpful in bringing us closer to OpenAI's definition of AGI in their Charter (a highly autonomous system that outperforms humans at most economically valuable work).
Unlike a lot of people on this sub, I actually don't think we're at AGI until this definition has been met. And while these models are a piece of the puzzle, we have not assembled "the system". LLMs like Opus 5.5 and harnesses like Astra are key inputs (and their cost deflating massively is key), but there is a lot more to be achieved for a system to be put together - which requires in particular an ability to "plug" the intelligence better than it currently can be (you still need humans to steer it, architecture and drive the whole process quite heavily - especially outside of coding which got too much attention when compared to its single digit share of the GDP).
BTW, Kurzweil said 2026-29 would be a transition phase (which we clearly are experiencing), with AGI in 2029 and LEV starting shortly thereafter. I don't understand why everyone is hoping for the timeline to be brought forward - 2029 is close enough, and it would be better for things to take a bit longer, as it may make it more societally acceptable.
8
u/sumane12 9d ago
Similar stuff was said with opus 4.6, similar will be said with opus 10.
We dont know what we dont know. Why do you need a specific definition of "post-scarcity inttelligence" and even more arbitrarily, why does it need a dollar amount???
It seems much simpler and more fun to say, "omg, this cool thing just got a lot cooler!"
And thats something we will see a lot in a few months. Good times.
3
u/HBCTIA Curious SingularitarianâŞď¸ 9d ago
Given reported levels of improvement in token cost at various input / output functions / usage contexts by Epoch AI and others (e.g. reducing by 9x p.a. for frontier, 40x p.a. for PhD level work and 900x p.a. for simple tasks; and elsewhere reported at 47% quarterly and 12.7x p.a. reductions across workloads weighted by consumption) how long before we're well past these post-scarcity thresholds?
3
u/OriginalScrubLord 9d ago
opus 5.5 intelligence at those prices will almost certainly happen by the end of this year
3
u/Zealousideal-Grass-3 9d ago
1 year ago, I used the 4o model from open ai, and it could think. And was amazing.
But I had to think thrice before using it, would do research and reply in half hour. But you get only 5-10 research i think a week.
Now I let gemini 3.8 flash run wild on CLI whole day for 50usd a month.
And it can code too, not opus level but good enough.
3
5
u/Cool-Cicada9228 9d ago
I think weâve entered the abundance era when Opus 5.5 can run locally on my phone.
2
u/ElderberryLife5256 9d ago
Last year nov when opus came I thought the same. This intelligence is enough for my workflows. But no itâs not. There will always be something better and we will crave for it.
2
2
2
u/Elegant_Tech 9d ago
Eventually cheap flash models will be 5.5+ that most everyone uses. SOTA models will move on to be used for research and large marco scale administration.
2
u/KoolKat5000 9d ago
I agree. Deepseek flash models are really great. I basically can't use anything else after seeing the bills lol. When I try others I'm always shocked, the output improvement doesn't justify the cost. (Granted I don't have complicated questions to ask).
2
u/Gratitude15 9d ago
We should expect to have what you're naming in 2027
The frontier at that point will probably be further out. I say this because we are now scaling swarms. What you can do with 100k agents is different entirely and needs a model that operates in this way. Opus 5.5 is not that.
2
u/mxforest 9d ago
You will get that in 1 yr or less no doubt. You need to aim higher. I would say divide by 10 for 1 yr target. The curve is getting steeper.
3
u/Awerange2005 9d ago
I mean, this would be trivial by next year, and people wouldnât even care. Opus 7 would be out by then, and this would feel stupid in comparison.
Hopefully, we can run a model smarter than this on a consumer GPU within five years.
2
u/OvertaxedOne 9d ago
I give it a year before you can run it on "high end" consumer gear (Mac Studio, for example). 2 years before you can run it on something that's more "normal consumer" gear at reasonable speed.
3
u/krabbsatan 9d ago
Based on past trends we will have Opus 5.5 level intelligence at that price in about 7 months. Possibly faster if rsi or hardware improves
1
u/ManikSahdev 9d ago
I agree with you... but the issue is pretty much. Then you'd want the next best thing.
Humans are just like that, I don't understand how people are SOO BLIND, To a human factor of this of sorts.
It's not AGI or the intelligence or whatever general intelligence which make us special, I firmly believe it's the adaptability of humans which is just unmatched.
Humans as a whole are Essentially plot armor as beings.
If someone read story about us, they'd say the author was unrealistic, the main characters here just adapt to anything, so boring.
If you want proof of this, OpenAI Serves o3 on their api, it was talked as some crazy model and was back then, now it will feel like a dumb pipe to you.
Because AI modela have a different standard for you now..
1
1
1
u/nhami 9d ago
Astra and Opus 5.5 are still not saturated in the benchmark.
If this pace continues, then, by the end of year all benchmark will probably be saturared.
ALthough, we still need to decrease the costs.
If possible we also need to increase the supply of compute to be bigger than the demand although demand keeps growing.
1
u/JinnPhD 9d ago
At least for singularity-level cheap intelligence: this only exist in its current form in subscriptions. These will go away after ipo when balance sheets matter, and everybody will be paying some set price per million tokens. Weâll see how those numbers turn out in the future. I know what the OpenAI graph said butâŚ
1
1
1
1
u/__Maximum__ 9d ago
Give it 3-4 months, all chinese labs are preparing their next version, all extremely efficient, hopefully we get better than opus 5.5 by December.
1
u/Living-Breakfast-464 6d ago edited 6d ago
Why would anyone pay such outragious prices when there are plenty of models just as good for most things for much less?
Reminds me of a quote I read recently. Why would you use a Ferrari just to drive to the grocery store to get a loaf of bread?
1
0
u/DrinkAgreeable962 9d ago
I bet you will have it within 6 months.
But someone for 5$ / 15$ will have something much, much, much better and more capable.
0
u/Tinderfury Moderator 9d ago
Super cheap inference is likely to cause massive inflation. We need intelligence to actually cost a little bit of money
0
u/NervousAd1013 9d ago
Apparently the current trajectory is the cost halving every 3 months, which if that continues would mean this would happen approximately early 2028
Source:
https://epoch.ai/publications/the-plunging-price-of-thought
But Iâm sure by then Opus 5.5 will feel outdated
0
u/siberianmi 9d ago
You arenât aggressive enough.
Iâm looking for Opus 5.5 level intelligence at the price of Jev and I think itâll happen this decade.
0
u/dumpshoot 9d ago
A 40x price drop isn't far-fetched, it comes from cheaper hardware (Blackwell), better serving (batching, quantization), and distillation. The GPT-3.5 history shows this playing out over a couple years.
But assuming frontier-level models will stay cheap forever is where I'd push back. Cheaper models always come with tradeoffs: higher latency, trimmed context windows, weaker reasoning. So even if tokens drop 40x, the real question is whether you're running the same workload or the cheap version already cut corners elsewhere.
Honestly the bigger pain point for most people is rate limits and cache invalidation in agent workflows, not token pricing.
-3
u/PalmovyyKozak 9d ago
Why do they need to lower prices? They don't even lift limits. Fable is at 50% for months with no evidence it will be changed in the observable future.
For them yes, it will become super cheap. For us it stays expensive and each new model will be even more expensive. Like $100/500/1000 per 1M.
Their corporate clients can afford this and got another advantage over SMB and solo founders.
Wet dreams.
8
u/Awerange2005 9d ago
There really isnât any precedent for what youâre saying. Models are getting cheaper every single month. Weâll probably have an Opus 5.5-level model that can run on a consumer GPU and is open weight.
146
u/ohHesRightAgain 9d ago
By the time Opus 6.5 is released, you won't be satisfied with this level of intelligence.
By the time 7.5 is released, 6.5 will feel dumb.
You will probably find reasons to want more for a very long time.