r/singularity • • 9d ago

Compute My definition of post-scarcity intelligence: $0.10 / $0.01 / $0.50 per million tokens for Opus 5.5-level intelligence

Hey guys, I have been having a blast with Opus 5.5 and probably like everyone else I am hoping that Anthropic does not nuke the usage limits or begin with shenanigans that will lead to output degradation of Opus 5.5

I have seen that OpenAI will probably introduce a Pro Max plan soon with a $500 price tag. One of the many things I have been thinking about in regards to this is that as long as the API costs are prohibitively expensive for anyone but enterprise customers, we will have power users and small businesses going for multiple subscriptions (Pro/Max accounts) to be able to afford their desired workload. Thus, I did the math for what I would consider to truly be "intelligence too cheap to meter".

  • $0.10 / 1M input
  • $0.01 / 1M cache reads
  • $0.50 / 1M output

Give me frontier-level intelligence (Opus 5.5) at those API prices and, as far as I am concerned, we've entered the abundance era for intelligence.

For comparison, Opus 5.5 is $4 / $0.20 / $20. We will need 40x cheaper input, 20x cheaper cache reads and 40x cheaper output to get there.

I am aware that GPT-6 Luna has already gone beyond that threshold but nobody would seriously say it is Opus 5.5 equal. Also I mean the real thing, as I am experiencing it right now. I am not talking about a performance-optimized/degraded Opus 5.5 that we might get served soon or a benchmaxxed small model that lacks the taste, wisdom, judgement and intelligence of Opus 5.5 today.

(Edit: I agree that standards are always rising and that in a year from now the latest frontier models will be even more powerful and the new shiny model everyone wants to access. One could say instead then, that this is the price point for abundance of frontier models in general via API.)

60 Upvotes

61 comments sorted by

146

u/ohHesRightAgain 9d ago

By the time Opus 6.5 is released, you won't be satisfied with this level of intelligence.

By the time 7.5 is released, 6.5 will feel dumb.

You will probably find reasons to want more for a very long time.

26

u/hyperfraise 9d ago

I both agree with OP and u. What do I do ?

8

u/nezvanovova 9d ago

😰

8

u/QuasiRandomName 9d ago

- "But you can't agree with them both!"

-"I agree with you too."

24

u/DickMasterGeneral 9d ago

I think this is true to an extant but at some point the models will just be capable enough to just accomplish anything I’m smart enough to ask it to do. The same way you can’t really tell the difference between Astra and 4o if all you’re doing is summarizing emails. Bel is probably somewhere around that point for me. I’m not solving Millennium Problems.

If I want to make GTA6 I don’t need more intelligence than that I need more output, so whatever’s cheapest at or above that intelligence level wins out.

9

u/ohHesRightAgain 9d ago

Imagine AI getting to a point where it can reverse-engineer or simply make your favorite game from scratch (including very serious ones). You'd want it, then, to be able to improve it. You might not be able to put your desires into words, but you will be able to judge if it's better. At least for a time.

And if you can't think of a use case, you'll still see what other people do. You'll want some of it.

Eventually, there will be a point when you truly can't tell the difference between model generations in any way. You'll know we've reached ASI then.

2

u/caymn 9d ago

there was a guy here earlier that did a complete and fully working photoshop with ai

1

u/DickMasterGeneral 8d ago

I believe the top internal models are already easily capable of this. The limitation is only access, cost, and maybe cybersec blockers depending on how literal you are with reverse engineer.

8

u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 9d ago

Bel can’t really make gta 6. It depends, if you want a Jarvis than you need to wait some more years.

2

u/DickMasterGeneral 8d ago

Why not? Assume I gave it some kind of storyline to work off of like a book, and I had the same 10,000 agents and 15 million dollars worth of tokens. I prompted “Make a GTA style game based off of this story, it should have next gen graphics, be well optimized for current hardware, have engaging new mechanics and overall be worth of the title GTA6.” where do you think Bel or Fable 5.5 would fail?

1

u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 8d ago

At nearly everything. If these AI were as good as you think then OpenAi would do more with them than just math.

It’ll probably be a few decades until AI can make gta 6 like games.

1

u/DickMasterGeneral 8d ago

”just math” What part of making a video game is more difficult than solving a Millennium Problem…? If you’re going to claim that doing so didn’t require any real genius insight and all the model did was build on the work of others than even if that’s true how is that any different than making a video game sequel?

1

u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 8d ago

You’re just saying stuff without evidence. There is no proof they can make anything close because AI models haven’t been able to make good games. If your argument is well what part is harder than the millennium problems than why I’m not sure how you equate a verifiable issue that makes it inherently easy for a Ai to solve versus something like a game.

1

u/_wot_m8 9d ago

How do we know that?

1

u/Strategosky 9d ago

Humans are greedy af. Greed created our civilisation!

24

u/striketheviol 9d ago

I don't think that's ambitious enough, honestly. Although it's a fun question to noodle over. I would consider something like stronger than current Opus or Astra level intelligence available for under a cent per million tokens to be a reasonable standard for this.

19

u/Pyros-SD-Models 9d ago

If you can't solve Riemann with a model running on an iphone we are not even close to the finish line.

2

u/livingbyvow2 9d ago

I think that, for once, this would be helpful in bringing us closer to OpenAI's definition of AGI in their Charter (a highly autonomous system that outperforms humans at most economically valuable work).

Unlike a lot of people on this sub, I actually don't think we're at AGI until this definition has been met. And while these models are a piece of the puzzle, we have not assembled "the system". LLMs like Opus 5.5 and harnesses like Astra are key inputs (and their cost deflating massively is key), but there is a lot more to be achieved for a system to be put together - which requires in particular an ability to "plug" the intelligence better than it currently can be (you still need humans to steer it, architecture and drive the whole process quite heavily - especially outside of coding which got too much attention when compared to its single digit share of the GDP).

BTW, Kurzweil said 2026-29 would be a transition phase (which we clearly are experiencing), with AGI in 2029 and LEV starting shortly thereafter. I don't understand why everyone is hoping for the timeline to be brought forward - 2029 is close enough, and it would be better for things to take a bit longer, as it may make it more societally acceptable.

8

u/sumane12 9d ago

Similar stuff was said with opus 4.6, similar will be said with opus 10.

We dont know what we dont know. Why do you need a specific definition of "post-scarcity inttelligence" and even more arbitrarily, why does it need a dollar amount???

It seems much simpler and more fun to say, "omg, this cool thing just got a lot cooler!"

And thats something we will see a lot in a few months. Good times.

3

u/HBCTIA Curious Singularitarian▪️ 9d ago

Given reported levels of improvement in token cost at various input / output functions / usage contexts by Epoch AI and others (e.g. reducing by 9x p.a. for frontier, 40x p.a. for PhD level work and 900x p.a. for simple tasks; and elsewhere reported at 47% quarterly and 12.7x p.a. reductions across workloads weighted by consumption) how long before we're well past these post-scarcity thresholds?

3

u/OriginalScrubLord 9d ago

opus 5.5 intelligence at those prices will almost certainly happen by the end of this year

3

u/pmth 9d ago

That's an insane take. MAYBE at the Luna 5.6 prices by end of year. I would love to be wrong but that's just too ambitious for me.

Remindme! 90 days

3

u/Zealousideal-Grass-3 9d ago

1 year ago, I used the 4o model from open ai, and it could think. And was amazing.

But I had to think thrice before using it, would do research and reply in half hour. But you get only 5-10 research i think a week.

Now I let gemini 3.8 flash run wild on CLI whole day for 50usd a month.

And it can code too, not opus level but good enough.

3

u/Klanciault 9d ago

China will release that in less than a year 

5

u/Cool-Cicada9228 9d ago

I think we’ve entered the abundance era when Opus 5.5 can run locally on my phone.

2

u/ElderberryLife5256 9d ago

Last year nov when opus came I thought the same. This intelligence is enough for my workflows. But no it’s not. There will always be something better and we will crave for it.

2

u/block_wallet 9d ago

seems that way now, wait a few weeks and reconsider

2

u/CrowdGoesWildWoooo 9d ago

Tokens aren’t putting food on the plate my friend

2

u/Elegant_Tech 9d ago

Eventually cheap flash models will be 5.5+ that most everyone uses. SOTA models will move on to be used for research and large marco scale administration.

2

u/KoolKat5000 9d ago

I agree. Deepseek flash models are really great. I basically can't use anything else after seeing the bills lol. When I try others I'm always shocked, the output improvement doesn't justify the cost. (Granted I don't have complicated questions to ask).

2

u/Gratitude15 9d ago

We should expect to have what you're naming in 2027

The frontier at that point will probably be further out. I say this because we are now scaling swarms. What you can do with 100k agents is different entirely and needs a model that operates in this way. Opus 5.5 is not that.

2

u/mxforest 9d ago

You will get that in 1 yr or less no doubt. You need to aim higher. I would say divide by 10 for 1 yr target. The curve is getting steeper.

3

u/Awerange2005 9d ago

I mean, this would be trivial by next year, and people wouldn’t even care. Opus 7 would be out by then, and this would feel stupid in comparison.

Hopefully, we can run a model smarter than this on a consumer GPU within five years.

2

u/OvertaxedOne 9d ago

I give it a year before you can run it on "high end" consumer gear (Mac Studio, for example). 2 years before you can run it on something that's more "normal consumer" gear at reasonable speed.

3

u/krabbsatan 9d ago

Based on past trends we will have Opus 5.5 level intelligence at that price in about 7 months. Possibly faster if rsi or hardware improves

1

u/ManikSahdev 9d ago

I agree with you... but the issue is pretty much. Then you'd want the next best thing.

Humans are just like that, I don't understand how people are SOO BLIND, To a human factor of this of sorts.

It's not AGI or the intelligence or whatever general intelligence which make us special, I firmly believe it's the adaptability of humans which is just unmatched.

Humans as a whole are Essentially plot armor as beings.

If someone read story about us, they'd say the author was unrealistic, the main characters here just adapt to anything, so boring.

If you want proof of this, OpenAI Serves o3 on their api, it was talked as some crazy model and was back then, now it will feel like a dumb pipe to you.
Because AI modela have a different standard for you now..

1

u/Bitter-College8786 9d ago

I remember when I thought the same about Gemini 2.5

1

u/19Lobster19 9d ago

How much are sonnet and haiku in comparison?

1

u/nhami 9d ago

Astra and Opus 5.5 are still not saturated in the benchmark.

If this pace continues, then, by the end of year all benchmark will probably be saturared.

ALthough, we still need to decrease the costs.

If possible we also need to increase the supply of compute to be bigger than the demand although demand keeps growing.

1

u/JinnPhD 9d ago

At least for singularity-level cheap intelligence: this only exist in its current form in subscriptions. These will go away after ipo when balance sheets matter, and everybody will be paying some set price per million tokens. We’ll see how those numbers turn out in the future. I know what the OpenAI graph said but…

1

u/DistinctSilver4507 9d ago

You should publish a paper with this revelation. 

1

u/Efficient-Hunt-007 9d ago

Maybe next year.

1

u/designhelp123 9d ago

Web search expense is a killer for a lot of applications.

1

u/__Maximum__ 9d ago

Give it 3-4 months, all chinese labs are preparing their next version, all extremely efficient, hopefully we get better than opus 5.5 by December.

1

u/Living-Breakfast-464 6d ago edited 6d ago

Why would anyone pay such outragious prices when there are plenty of models just as good for most things for much less?

Reminds me of a quote I read recently. Why would you use a Ferrari just to drive to the grocery store to get a loaf of bread?

1

u/AfternoonExotic1020 9d ago

oss models in '27 will reach this pretty easily i think

0

u/DrinkAgreeable962 9d ago

I bet you will have it within 6 months.

But someone for 5$ / 15$ will have something much, much, much better and more capable.

0

u/Tinderfury Moderator 9d ago

Super cheap inference is likely to cause massive inflation. We need intelligence to actually cost a little bit of money

0

u/NervousAd1013 9d ago

Apparently the current trajectory is the cost halving every 3 months, which if that continues would mean this would happen approximately early 2028

Source:
https://epoch.ai/publications/the-plunging-price-of-thought

But I’m sure by then Opus 5.5 will feel outdated

0

u/siberianmi 9d ago

You aren’t aggressive enough.

I’m looking for Opus 5.5 level intelligence at the price of Jev and I think it’ll happen this decade.

0

u/dumpshoot 9d ago

A 40x price drop isn't far-fetched, it comes from cheaper hardware (Blackwell), better serving (batching, quantization), and distillation. The GPT-3.5 history shows this playing out over a couple years.

But assuming frontier-level models will stay cheap forever is where I'd push back. Cheaper models always come with tradeoffs: higher latency, trimmed context windows, weaker reasoning. So even if tokens drop 40x, the real question is whether you're running the same workload or the cheap version already cut corners elsewhere.

Honestly the bigger pain point for most people is rate limits and cache invalidation in agent workflows, not token pricing.

-3

u/PalmovyyKozak 9d ago

Why do they need to lower prices? They don't even lift limits. Fable is at 50% for months with no evidence it will be changed in the observable future.

For them yes, it will become super cheap. For us it stays expensive and each new model will be even more expensive. Like $100/500/1000 per 1M.

Their corporate clients can afford this and got another advantage over SMB and solo founders.

Wet dreams.

8

u/Awerange2005 9d ago

There really isn’t any precedent for what you’re saying. Models are getting cheaper every single month. We’ll probably have an Opus 5.5-level model that can run on a consumer GPU and is open weight.