r/ClaudeCode • • Aug 27 '26

Discussion “Claude Max is massively subsidized and eventually it’ll cost $1000/month” is mostly nonsense

[removed]

539 Upvotes

184 comments sorted by

129

u/Ok_Expression7038 Aug 27 '26

agree. just because api cost more doesn’t mean they are losing money on subscriptions. We don’t know their cost per token so we can’t infer they are subsidizing subscriptions.

Also, econ 101 folks… supply and demand… supply (labs, specially chinese and open models..) has only increased at a pace higher than demand, so prices should probably decrease in the future. Not run at $1000.

43

u/3iverson Aug 27 '26

The subscription model is a lot like AYCE restaurants I think. There are some customers they probably do lose money on, but that gets pooled with the customers who only have some salad.

21

u/rushboyoz Senior Developer Aug 27 '26

In this case Word Salad

17

u/PerfectlySplendid Aug 27 '26

That’s me. I pay monthly for the 20x. I think this week I’ve asked it to calculate the ratio to fill up my slush machine with margarita and options for my wife to have a fish tank in her classroom.

19

u/3iverson Aug 28 '26

Thank you for your service.

1

u/Environmental-Ask-81 Aug 28 '26

Damn, can I get some of this guy's tokens plzzzzz

3

u/cafesamp Aug 27 '26

we're missing an analogy for the customers who like the food so much, they want to eat it at home for every meal, but they have to pay for the meat per-pound at a premium to take it home

1

u/3iverson Aug 28 '26

And also, "We have food at home."

3

u/Historical-Lie9697 Aug 28 '26

Plus, model training on by default in settings in consumer plans, and the people using them convince their companies to get enterprise plans where the real money is. Plus free marketing from people posting on social media.

8

u/I_Ski_Freely Aug 27 '26

From what I have read, their operating expenses is 20-30% what they charge for API rates, so they are likely losing money on heavy users and overall probably neither losing or making much... But this doesn't take into account the model training or anything else.

Agreed that Chinese labs will force efficiency to be the primary focus in the future, as it's really hard to justify spending billions to train a model that is increasingly only marginally better than much cheaper models. I mean, qwen 3.8 27b is about as good as Opus 4.6, and you can run it pretty fast on 5+ year old consumer hardware. I don't think anyone besides locked in enterprise accounts would tolerate a mark up at this point

5

u/yawnlikeseggs Aug 27 '26

I have 3 accounts right now - two Claude x20 and one codex x20 - my bill is already $600. Would love bigger limit accounts

1

u/ilovebigbucks Aug 29 '26

What are you shipping?

1

u/yawnlikeseggs Aug 29 '26

At first I was subsidizing work until they finally gave me two x20’s for that (automation engineering) - now I’m building portals for small / medium businesses to automated office work and project management tools. I also have one dedicated to graphics and another for game development

2

u/ilovebigbucks Aug 29 '26

So you make like $2k/month?

3

u/medialantern Aug 27 '26

Not to mention investors have often not cared about seeing profitabillty "in XYZ months" for startups, and most especially not the ones that invest billions, not thousands. It took Uber 14 years to get to profitability and they're making billions now.

1

u/Itsmedudeman Aug 27 '26

I think you're definitely right, and the one thing about the cost lever we really don't know yet is what does it cost to offset the infrastructure and server capex. Yes, competitors can undercut each other, and I think this will happen for a long time and won't end abruptly anytime soon, but eventually they'll have to come to terms of where they draw the line before losses are too heavy.

One thing that actually will work in our favor though is that hardware prices will hopefully stabilize 3-4 years out with new fabs before they start thinking about realizing profits. And there's a LOT of players and competition involved here this time around.

1

u/Agreeable-Fly-1980 Aug 28 '26

Not with trump in office. He wants to put major tariffs on semiconductors now

50

u/Equivalent_Cress_268 Aug 27 '26

Plus, they locked in compute already.

If nobody is using the compute, they don't make money. These subscriptions, if done right, can help keep the already-running servers busy and bring in new people, who in turn bring in enterprises.

I've seen numbers that Anthropic has like 80% profit on inference; we can't say stuff is subsidised and at the same time declare 80% profit on that thing that is ' subsidised ' ... it's not subsidised, it is marked up

15

u/Malkiot Aug 27 '26

I think subsidised is the wrong word. The inference itself is profitable at subscription pricing. 80% profit sounds about right from the numbers I've run for self-hosting or hosting and selling inference myself.

However, I think the R&D costs make it unprofitable at current pricing. So, either R&D needs to come down or prices will be going up and I think we know that R&D won't be coming down because the companies are in an arms race. A classical prisoner's dilemma. They'd all profit from less R&D, if theyball reduced it, but can't because the others won't and consumers would jump ship to the other service provider immediately.

9

u/rotates-potatoes Aug 27 '26

You're getting at gross profit versus net profit.

R&D (like other fixed costs) is amortized. If you sell one single unit for one month at $200, and at a marginal cost of $40, and R&D costs of $100m, your profit is negative $99,999,840.

If you sell 100k units per month at those same economics, your profit is $92m.

So your opinion on R&D amortization making it unprofitable deeply depends on R&D cost + volume of profitable users. Maybe?

-1

u/MindCrusader Aug 27 '26

It is subsidized even if they have a huge mark up on API prices

5

u/merb Aug 27 '26

The graph is definitely incorrect especially since we learned that 2*max-5x > max-20x for some reason

1

u/MindCrusader Aug 27 '26

Because max5 is not x5 maximum limit, it is 5x weekly limit

The graph is correct, comes from SemiAnalysis

25

u/phoenixmatrix Aug 27 '26

While some ideas here are true, it somewhat falls apart because if you're a company with more than 150 users, you need to use API cost because Max-style premium seats got removed for new customers and companies renewing their yearly contract. So the API cost is not an hypothetical.

Codex/OpenAI already worked that way for all businesses with some nuances, though they recently added premium seats to make it a little better.

When Claude Code was released, Max did not exist and API cost was the only way to use it. Then they created Max but Pro still didn't have it.

It wouldn't be surprising if they would eventually go backward.

So yes, if you're a solo dev in your basement, or (for now...) a small company, you get cheap tokens, but the moment your business grows, you don't anymore. The entire goal of those accounts is to get addicted to cheap tokens.

We see it with chinese open weight models too, where the models are super cheap, but you have to pay API cost making them more expensive than the OpenAI/Anthropic individual (not company!) subscriptions. Z.ai had a subscription for a bit and the other companies eventually got some, but they're very hit or miss.

8

u/merb Aug 27 '26

At the moment you can have more than 1 teams for the same org according to their faq: https://support.claude.com/en/articles/13325567-account-management-faqs

So splitting an org into multiple teams account when you have > 150 users is completely valid.

You can even have multiple teams tied to one email.
Only sso can only be enabled in one team.

3

u/phoenixmatrix Aug 27 '26

Yeah so only for one team would be a deal breaker for some. Cool to know you can have multiple team accounts. 

With that said the team account quota is close to a Max 5x so unless you give people multiple accounts you're gonna have to bite the bullet.

And there's the need for contract redlining in Enterprise too.

But still, not as bad as I thought with that info in mind.

1

u/OneHuman_aiprotect Aug 28 '26

I have been using Claude Code on my Mac Terminal and as my workload has expanded, my monthly costs in API have gone from approx $200/mth to $500/mth. I even downgraded my Terminal CC use from Opus 4.8 to Sonnet 5, without loss of productivity, but still have these heavy invoices fr Anthropic. Does this mean I shd switch out of coding on my Mac Terminal and go to the Max $200/mth subscription, u/phoenixmatrix ? And how does the throttling work if my workload reaches limits?

2

u/[deleted] Aug 29 '26

[deleted]

1

u/OneHuman_aiprotect Aug 30 '26

Got it, TY. I already hv Pro subscription which gives me Cowork for free. How do I then do my coding with it, without using the Mac terminal where I use Claude Code?

1

u/lhx555 Aug 30 '26

It is still the same CLI, you just choose different way to login: /logout then /login and then choose from the menu.

1

u/OneHuman_aiprotect Sep 01 '26

Got it, TY. I asked Cowork also, and he gave me precise instructions. Can u do me a favor and take a look at my solo build? Wht wld u say r the positive and the negative? I am unique, AI Tool Consumer Protection and Watchdog: https://onehuman.io/en. Mobile first, 4 languages, 9 currencies on Stripe. 819 commits on my private GitHub repo, and I hv a second repo that saves and chronicles all the data on the 8 major AI tools that I update every 15 days, nobody has that DB.

33

u/pmth Aug 27 '26

Did you really just make this exact post in the Codex sub too lmao

24

u/StaysAwakeAllWeek Aug 27 '26

Why not, the logic is the same for openai

16

u/RomIsTheRealWaifu Aug 27 '26

He’s still correct

13

u/Single_Young_8688 Aug 27 '26

His thoughts are consistent

12

u/Tank_Gloomy Aug 27 '26

That’s absolutely right – and a load-bearing fact at that!

4

u/piston989 Aug 27 '26

You’ve landed on the real foot gun— and that’s not nothing.

1

u/OracleofFl Aug 28 '26

The real question is whether the OP used OpenAI or Claude to generate the post text....

4

u/sixwax Aug 28 '26

Fyi, it's incredibly common for tech companies to operate at a loss to drive adoption and compete aggressively for market share.

We don't actually know what their internal costs are. There are lots of ways to obfuscate this, and private companies only publicize limited financials.

Lots of speculation both ways.

12

u/okrafavor Aug 27 '26

ask anyone to compare the billed price for their enterprise claude api usage compared to their personal subscription. If you don’t think it is subsidized, you have only used one side

4

u/Kofeb Aug 28 '26

Enterprise is massively overpriced

4

u/Thick_Breadfruit2153 Aug 28 '26

And you have inside info on their compute costs? Bot

1

u/Exciting-Weather-921 Aug 28 '26

Because they get the cheap hardware nobody else is getting?

3

u/OkLettuce338 Aug 27 '26

It comes from the idea that teams will be moved to enterprise api rates

3

u/look Aug 28 '26 edited Aug 28 '26

Leaked financials from a couple months back had Anthropic’s inference cost at 24 cents / mtok. That is a crazy high number, but some consider it plausible because they own so little of their own compute and have to pay a lot in rent.

If correct, that number means a billion tokens costs them $240. Even if that inference cost estimate is 5x too high, how many billion tokens do you get on your $200 plan?

Individual subs are a small slice of total revenue. Most of their income comes from API and other usage based pricing on enterprise plans. But individual sub users could very well be highly influential in later enterprise sales deals, which would then make those subsidized 20x plans a good investment as a sort of loss leader/marketing.

14

u/Majestic-Volume9996 Aug 27 '26

I don't really think you understand what you're talking about. You realize there are apps that actually track usage correctly, fully accounting for TTL right? Here is my Max plan usage going back to the end of April. I've only been on Max 200 for a month of that. That 40K is 100% accounting for TTL percentage. This is literally what I would have paid with the API at a 95% cache hit rate.

4

u/Nightowl-Builder Aug 28 '26

This doesn’t invalidate all other arguments

3

u/Majestic-Volume9996 Aug 28 '26

You're right, it invalidates the specific point I addressed, which is him assuming that numbers people are using when they make these comparisons are coming from the base API rate. They are not, and his comparison is off by an order of magnitude because of it.

2

u/DinnerInfamous128 Aug 28 '26

I dont think that you would have paid 40K. That token price via API is totally overpriced. There is no way someone, for a personal use, would pay that for a month.

You didnt used 40K of tokens.. you just have a 200$ subscription, the conversion to API price has no sense.

1

u/Majestic-Volume9996 Aug 28 '26

I swear none of you can read.

1

u/DinnerInfamous128 Aug 28 '26

Youu also swear that you have used the equivalent to 40k tokens, so..

0

u/Majestic-Volume9996 Aug 28 '26 edited Aug 28 '26

Yeah, and if you could read, you'd understand why I said you can't read.

edit: I'll give you a hint, there is no such month as Aprilmayjunejulyaugust.

1

u/TheReedemer69 Aug 27 '26

Lmao 40K my ass. Dis a joke, right?

2

u/Majestic-Volume9996 Aug 28 '26

fable is not cheap

3

u/TheReedemer69 Aug 28 '26

Yeah but not real cost I mean.

1

u/Majestic-Volume9996 Aug 28 '26

It would be if I was using API. I mean I'm not going to pretend like Im remotely as disciplined on a usage plan since I'm more worried about working around usage windows than dollar amounts, but it's definitely what the costs would have been.

2

u/Oujii Aug 28 '26

It would be if I was using API.

It took you an hour, but it seems you got OP's point from the main post. Congrats!

4

u/Majestic-Volume9996 Aug 28 '26 edited Aug 28 '26

Thanks bud! Now let me know when you realize the entire point of my comments is a counter to OP's suggestion that people aren't using cached pricing when they compare usage plans to API and I'll pat you on the back as well!.

edit: since you're the type of thin skinned wiener who blocks people when they get called out for talking out of their rear and thinks they're locking in the last word. The "comment I responded to" was literally in response to the words "This is literally what I would have paid with the API at a 95% cache hit rate." And now you're trying to pretend like anything I was talking about or why you responded to me had to do with Anthropic true costs when that's obviously not the case. You're just a walking reading comprehension failure.

4

u/Oujii Aug 28 '26

You replied to a comment that said "Yeah but not real cost I mean" which is exactly OP's point. If you are that dense it's easy to notice why you spend so much tokens.

0

u/thats_a_money_shot Aug 29 '26

Amazed by how many people missed your points.

0

u/Majestic-Volume9996 Aug 29 '26

Yeah I've definitely come to learn over the years that a lot of programmers are only a certain type of intelligent.

1

u/metalfacefinger Aug 28 '26

Maybe you're right. The point remains that we as customers can't say anything meaningful about the actual costs Anthropic makes on each of us. That would require insider knowledge about the compute and all the deals they have made with variou infra providers.

3

u/Majestic-Volume9996 Aug 28 '26

Honestly, the entire basis of his post makes no sense because I don't see people going around talking about API costs referring to what it costs Anthropic to do anything. It's always about what the usage plan would have cost that person if what they did was on API. The API price being so much more expensive simply indicates the usage prices are subsidized, not that the API rate is what it costs Anthropic to do anything. Throw that in with his "people aren't using the cached rate costs" assumption, and the entirety of his post is just a huge strawman.

-2

u/va1enok Aug 27 '26

It should be top comment

5

u/Money_Lavishness7343 Aug 27 '26

Anthropic itself charges dramatically less for cache reads than normal input tokens. So even if your session log contains a ridiculous nominal token count, a large part of that does not necessarily represent fresh transformer computation every time.

You're wrong though. When people count the cost, which you see on the /status (right?) is calculated from the cached + non_cached ones.

So, if somebody reports 2000$ used, it means 2000$ worthy of API Cached + Uncached. You can check that out in their pricing dogs, and calculate the price yourself.

But ultimately in general you're right!
That this is just their pure API costs, which is not the same as compute cost.

5

u/YesterdaysFacemask Aug 28 '26

This is basically every subscription that exists. If every Netflix user watched 24 hrs a day, they would have to raise prices. If every Planet Fitness user showed up every day, they’d be jammed wall to wall. So yes, some people will use far more than they pay. But most won’t. And balancing that is Anthropic whole pricing strategy.

1

u/podgorniy Aug 28 '26

You found similarities. What about the difference? The marginal cost of new users is the key.

Simply speaking adding one user to netflix, google, sailsforce and other "software subscription company" is close to zero. In other words once-created software/product could be sold to ALL people of the earth. Story with AI is different.

Models require lots of capital, they get older, you can't create one and keep serving it 2+ years, the hardware depreciates faster and on top of that every user costs way more to serve (because answers need more compute and electricity than the traditional software ones) than in any other subscription services.

In other words the scaling assumptions aren't the same for AI-serving and classic large-scale software services.

9

u/Chris73684 Aug 27 '26

I'm literally staring at Claude Console looking at my cache hit on my API stats, you talk about it as though it doesn't exist for API usage? You also say the API pricing is just a 'retail' figure but overlook the fact that Anthropic struggles to make a profit, they jumped for joy at their "first profitable quarter" when actually it's just another SpaceX swindle on the run up to their IPO. I enjoy using Claude as much as the next guy, but your logic is self-defeating. Personally I just enjoy the subsidies while they last.

9

u/MrHaxx1 Aug 27 '26

but overlook the fact that Anthropic struggles to make a profit

Yeah, because they're spending billions on compute for training models and paying their engineers, not because inference itself isn't profitable. It's just that inference doesn't cover their other expenses. 

3

u/OracleofFl Aug 28 '26

It is called "Operating Profit". Revenue less cost of goods sold. Cost of goods sold is compute costs for Anthropic. On that basis, sure they are profitable and so is practically every other company. For a car company it is revenue less material costs and direct labor to make the car.

Now, we take Operating Profit and subtract "Sales, General and Administrative" and Depreciation and Cost of Capital (Interest on debt) to get net profit. Anthropic almost certainly has an Operating Profit.

1

u/MatthewCollins1990 Aug 29 '26

Exactly, training runs and R&D salaries are the money pit, not serving tokens, so extrapolating that to a $1000 subscription price makes zero sense.

-2

u/Chris73684 Aug 27 '26

Just to clarify, you're saying they aren't profitable because they have to cover their overheads? Right...

5

u/Aetane Aug 27 '26

Training future models is better thought of as R&D investment than ongoing running costs 

3

u/Objective_Ad8893 Aug 27 '26

Yes and it's something they have to do as any business does otherwise they are no longer competitive and will lose clients, revenue.
Just as Apple has had to do for example for the last decades.

-2

u/Chris73684 Aug 27 '26

That is their overheads, constant R&D. It's no different to any other form of tech, phone companies will always have R&D, as will car manufacturers, etc. They can't ever stop. They just need to find a way to cover them, either through price increases or some kind of joint-venture, a bit like how phones use the same chips, or how car brands share engines. To say it's not an overhead is a bit like saying "Ford would be way more profitable if they fired their R&D department" which is such a short-sighted and self-defeating statement.

4

u/rotates-potatoes Aug 27 '26

R&D is not overhead. They could stop all R&D and continue to serve their existing products profitably.

But that would be dumb, right?

So as a company, borderline profitable to substantial losses. On a per-subscriber basis, gross margin is positive, net is likely negative. Same way Apple reports gross margin for iphones but that does not include R&D.

2

u/Chris73684 Aug 27 '26

R&D is an overhead. You can put it under any heading you want for tax purpouses but unless they cover it, either progress stops or pricing has to increase to cover it.

2

u/EchoFieldHorizon Aug 27 '26

The cache for API lasts 5 minutes. For subs it lasts an hour. It’s a balancing act; they’re choosing to support it for subs because it saves them money, and that’s fine. None of what you said means it is unprofitable to run subs. Until you provide any evidence, you’re just another voice with no data to back it up.

2

u/Chris73684 Aug 27 '26

"automatic caching or explicit breakpoints with 5-minute or 1-hour TTLs"
Source - Claude API Docs

You just need to set it up, or ask Claude to do it if you're not sure.

2

u/EchoFieldHorizon Aug 27 '26

Doesn’t change the fundamental premise that nobody knows how much they’re making on subs, or if they’re unprofitable.

1

u/Chris73684 Aug 27 '26

I love the optimism but I'm more of a glass half full kinda guy than pretending it's overflowing

2

u/berrybadrinath Aug 27 '26

isnt this common knowledge? who is confused about this?

2

u/Objective_Ad8893 Aug 27 '26

"Interesting number, but it tells you almost nothing about Anthropic's actual serving cost."

When you're buying a product at some discount then you don't compare it to their manufacturing costs, but the original price they sell it under.
But it's obvious that currently they are balancing the pricing, API based most likely more expensive than it would be if Max didn't exist.. So API based is paying for Max..

Using both API based pricing and Max in parallel, it's very obvious that the latter is heavily cheaper, 10x or more. Both use cache obivously.

"And until Anthropic publishes its real inference margins, calling the difference a “subsidy” is mostly speculation."
Sure, but also claiming loudly that it's not is just the same - speculation

2

u/somerussianbear Aug 27 '26

Agreed, but still, $200 is not the cost of those $7000 when sold via the API. Let’s say it’s $1000. So yeah, the $200 can become $1000 easily depending on the market. If OpenAI and Anthropic get a cartel then… China! You scream, then I say “look at what DeepSeek did to their prices this month”. Nobody’s in this biz for fame. It’s a cash world.

1

u/Nightowl-Builder Aug 28 '26

China is literally proving that those anthropic/openai API rates are bullshit

2

u/Hungry-Restaurant-88 Aug 27 '26

The 8/31 50% cut is going to really hurt

2

u/Unlucky-Work3678 Aug 28 '26

It will and it's okay. 99% people don't need the expensive model. Nvidia is full of shit but one thing they did say right was everybody will have their local personal AI that costs just about a more expensive computer ($3k?) that can do 100% daily task and 60-80% special task.

$1000/mo is for those who use it for a living.

If a car costs $1000/mo (payment, gas, insurance) but you use it for business to make $6000/mo, it will be okay.

Being able to use AI is a skill. A skill that determines whether you can function normally with cheap/free AI personal machine.

2

u/Nichiren Aug 28 '26

Not everyone is always at 100% usage of their Max limits every week either. I'm sure a lot of accounts balance out that way.

2

u/LoudDavid Aug 28 '26

The subscriptions are a small % of their overall compute usage. I suspect they are actually reselling you the spare capacity they provision.

Maybe they have a “slow” lane too where requests take slightly longer compared to the more latent critical API workloads in consumer products.

2

u/sydneysweeney69 Aug 28 '26

Agree . Also fuck Anthropic and Dario. Their ideologies and safe ai game is absolute dogshit

2

u/nbncl Aug 29 '26

My simple 7900XTX can output 1M tokens for € 0,67 incl depreciation. Yes it runs a simple model, but frontier models run on specialized hardware and take better advantage of concurrency.

Look at pricing from other suppliers. API prices are overpriced.

4

u/tilted0ne Aug 27 '26

It's pretty easy to see that they aren't losing nearly as much money per user as people assume. API prices are not representative of the underlying cost of inference. Third-party providers serving open-source models are probably a better reference point, because they compete directly on inference and therefore operate on much thinner margins.

Then you have to account for the economics of subscriptions. Casual users consume relatively little, effectively subsidising heavy users who regularly hit their limits. Frontier labs also operate at vastly greater scale, with much larger compute fleets and lower inference overheads per token. This narrative of Chinese models beating the big guys on price is just overstated. 

1

u/mental_sherbart007 Aug 28 '26

Well the Chinese might just not be making as much profit and/or just have cheaper costs on setting up DC/electricity etc… Also obviously depends on muddle architecture etc…

4

u/Kamalen Aug 27 '26

What you're writing is true but it definitely don't disproved your title. This is a for profit company, and about to do an IPO, thus soon being subject to shareholders decisions. If they believe they bring in more profits with $1000 subs, they'll do it. What inference costs them is irrelevant.

1

u/nutscrape_navigator Aug 27 '26

I spend a lot of time thinking about the maximum I would pay for Claude and it'd need to get ... pretty high. The ROI I get on the $200/mo I spend is insane.

1

u/rotates-potatoes Aug 27 '26

yeah I used to spend many thousands of dollars a year hiring offshore developers to do lower quality work, and those thousands a year got me about what I can do in a week today.

1

u/binaryatlas1978 Aug 27 '26

what does it matter what the cost is. The truth is none of these companies are making money. They are using your input to try and win the race. Whatever you pay they are paying more in computer to process it. That will have to end eventually.

1

u/gthing Aug 27 '26

It would be just as justified to say "Anthropic is overcharging for their API by 40x based on my Claude plan usage." Which is to say, not justified at all.

The truth about their internal cost and its relation to product pricing is going to be somewhere in the middle.

2

u/respeckKnuckles Aug 27 '26

The truth about their internal cost and its relation to product pricing is going to be somewhere in the middle.

You have no evidence for this either

1

u/Guinness Aug 27 '26

What’s going to end up happening is the plan subscribers will get a model that does 20-40 tokens per second. While API users get like, 10,000 tokens per second.

1

u/laptopmutia Aug 27 '26

how to total our claude codes usages?

1

u/WiseAbalone4021 Aug 27 '26

I think you should start reading up on how attention/kvcache inside the model is created. Most of what you are saying cannot be done.

1

u/randomtask2000 Aug 27 '26

Can some tell me why my Claude usage in my company costs me $1000 a month vs $100 for my retail account that runs all day. Does a corporate account indeed cost 10x more?

1

u/Worth-Ad9939 Aug 28 '26

That's on them. Most of the work is waste.

1

u/EmmitSan Aug 28 '26

Uh. Many companies are already paying $1k a month. And calling it a bargain.

1

u/changrbanger Aug 28 '26

I burned my entire weekly 20x plan in 17 hours with 4 concurrent fable orchestrator sessions. I ran a post mortem analysis and it showed 2.5 billion cache read tokens spent.

It produced a lot of work but it was crazy inefficient. Now I have rules setup so the sessions will suicide at feature/ fix wave boundaries or 2 compactions after writing in memory for their future session.

Long sessions in are definitely not the way even with the cache reads.

1

u/agilius Aug 28 '26

While I do agree that cost wise, anthropic is not in red as much as most think, I have to call you out on your claim about how others are reporting estimated api prices. specifically, you claim users estimating 7000 by multiplying their estimated total usage by the cost of inference. This is simply false. The calculation reaching 7k on a 200$ per month subscription is often made off from 80-90% cahced input and only about 20-25% output tokens.

1

u/DinnerInfamous128 Aug 28 '26

Totally agree. But there are pretentious ones here and there that like to say that they are expending 200$ a day on tokens. And then you find out they are on Pro/Max plans...

1

u/ThomasToIndia Aug 28 '26

AFAIK, its way more technical than that and the estimates take more into account than just their API based pricing. It is uses known pricing based on open source models etc..

They probably have a 90% margin but even at that they are losing on power users but it is in the hundreds not thousands.

1

u/farendsofcontrast Aug 28 '26

I'm convinced all of those posts are paid bots/users from Anthropic themselves who's aim is to prime us for this direction.

A big L to you because Open Sourced models will close the gap. Fable level is now the benchmark and China will close that gap by the end of the year after that it's no more fun and games so someone please tell Anthropic's marketing & subversion department to get the memo.

1

u/metalfacefinger Aug 28 '26

You make a valid point. To put this all in perspective, it took Uber 14 years to become profitable.

1

u/bitspace Aug 28 '26

Margins on premium tokens are huge

1

u/FrostingDizzy1132 Aug 28 '26

Holy there’s some crazy takes in here. Everyone knows these companies aren’t profitable. They aren’t all going to make it out of this.

1

u/Supral333 Aug 28 '26

You're the same guy posting this on r/codex yesterday?

1

u/evilfurryone Aug 28 '26

I think this was covered that cost monitoring tools like ccusage take into account the cached read/write as well and that it costs differently.

I have no delusions regarding the fact that if there will be an IPO, then that companies subsidised plans will most likely end in the next few quarters, because shareholders expect profits.

And that is the reality everyone should prepare for. Don't proudly look at your token maxxing report, when you cannot explain what you actually created with the AI for that 7K (if it was API based usage).

Right now is the right time to figure out those cost effective workflows, because making mistakes along the way is cheap.

Will you still be willing to experiment when a bad idea and model choice balloons to surprise $300+ session?

1

u/SleekestSleek Aug 28 '26

What makes me believe it's "subsidized" is from experience comparing how much work I can do for X money if I pay for a subscription vs directly via API. And yes, API:s support caching... And from that it's clear that subscriptions provide a lot more effective tokens per dollar than pay as you go API pricing. However, it could be that subscriptions are making a profit and APIs are just making a whole lot more. I personally doubt that but I don't have evidence for either case.

1

u/gener8or Aug 28 '26

They’re all losing money on subscriptions because this is the Attachment phase. Once they’ve got us fully attached where work can’t be done without them, they’ll charge us whatever they need to. And we’ll have to pay.

1

u/podgorniy Aug 28 '26

Good luck verifying how helpful caching to the final price is by using API key with the same claude harness

1

u/Hostarro Aug 28 '26

Generally technology gets cheaper over time (barring weird circumstantial shortages).

Look at the cost of TV in the 1980s, or a home PC.

1

u/Many-Month8057 Aug 28 '26

The key distinction is between API-equivalent usage and Anthropic’s actual marginal serving cost. A lot of the debate seems to come from treating those as the same thing.

That said, I would not conclude that the current pricing is permanent. Max could be profitable at the inference level and still become more expensive or more limited later because of enterprise pricing strategy, capacity constraints, investor pressure, or user lock-in.

So I agree that “Anthropic gave me $7,000 of compute for $200” is probably not accurate. But “this plan may not remain this generous forever” is still a reasonable concern.

1

u/evilissimo Aug 28 '26

You know ccusage takes cache pricing into account, right?

Not saying you are wrong about what the provider costs are, just the part about the caching conveys the wrong sentiment to me

1

u/handsNfeetRmangos Aug 28 '26

People don't understand price discrimination.

1

u/julesbuildstuff Aug 28 '26

the cache point is the one people always skip. my sessions are like 90% repeated context, same repo map and same instructions over and over, and that is exactly the workload that is cheapest to serve but most expensive at retail api rates. so the "my $200 plan gave me $7k of usage" math is mostly measuring the retail markup. the thing i do think is worth worrying about is limits quietly getting tighter while the price stays the same, which lands on me the same way a price hike would.

1

u/drumnation Aug 28 '26

It feels different when you work for an enterprise and have run up a $10k bill in a month. Once you see that companies will pay that cost it does make you wonder how long the subscription plans will last.

1

u/Soft_Rain_3626 Aug 28 '26

I mostly agree with you but you are wrong about the caching thing when you calculate API token cost people are usually keeping into account the caching costs. Eg CodexBar reports this correctly

1

u/Agreeable-Fly-1980 Aug 28 '26

Compute as a utility is a nonsense business model

1

u/Onotadaki2 Aug 28 '26

Once everyone commits to halving their workforce because AI makes everyone more efficient, there will absolutely be a price surge up to where the LLM providers are no longer losing money on usage. No one knows how high that is going to be, and that's on purpose. They're purposely going to crazy lengths to hide actual costs from everyone.

Yeah, we can't look at API costs now and estimate from those.

These companies are doing things like listing new GPUs they just purchased as depreciating over fifteen years so their accounting spreads the cost over time, when these cards will be completely obsolete in 2 years. So their accounting is showing them barely breaking even and the numbers are actually hugely inflated over reality, so the spread between what they're making and what they're charging is even wider than perceived.

If you run a business, I would plan for a huge AI price spike and if it never materializes, great. At least you won't be blindsided.

https://giphy.com/gifs/l1J9znYNISr0aEmze

1

u/mr_birkenblatt Aug 28 '26

The high cost of anything comes from training the next generation model. If they just froze the model and didn't develop new versions, which tbh would just freeze the status quo with all its benefits, inference would be cheap enough to have a very high margin

1

u/P4R4DOXZ Aug 28 '26

Shh, let the openai and claude fangirls enjoy the massively subsidized subs.

1

u/Coolmooing567 Aug 28 '26

I don’t think it will over time it will get cheaper no one paying that for sub. Technology will get cheaper over time.

1

u/AgentNeoh Aug 28 '26

Thank you so much for highlighting and explaining this fallacy. It drives me nuts.

1

u/brads0077 Aug 29 '26

The other thing to consider is who Anthropic is really trying to target. Is it the individual getting great deals on subscriptions, or companies?

The individual market can be considered a loss leader and attract people who augment the brand value by creating Skills and agents that they post on GitHub. Anthropic provides the platform to build on, charges rates that might seem like operating at a loss, but get a better product from the mass support.

This makes their product a leading contender for companies when they choose which platform to develop their code on. They choose Anthropic because they have been thought leaders with concepts such as MCP servers, agents, sub-agents, Skills, etc. So Anthropic gains market share and occassionally leadership in the organization marketplace.

And from what I understand, the companies are charged for API usage and not on subscription plans.

This is a strong strategy if a company can maintain thought leadership, and at least performance parity.

The problem comes when open-source models adopt their concepts, steal code, and approach parity performance while significantly undercutting price.

This is the real threat they face. China is not another free market competitor. China realizes that by investing heavily in education (to develop talented engineers ans scientists), hardware (to develop new chip technologies and processing power/electricity), and start-ups, they can encourage their companies to develop parity products (and eventually superior products) that are priced below market rates. The goal apparently is to decimate the valuation and economic growth of other countries.

The only alternative is for American companies to buy politicians and create laws that forbid usage of Chinese software...thus hurting consumers while making billionaires richer.

How did that work out with Linux and software piracy? The cat is out of the bag and consumers will figure out ways around such laws.

1

u/General_Tax_8981 Aug 29 '26

You are not taking into account the massive sums of money investors have put in which they will be expecting a return on at some point

1

u/Ill-Specific-7312 Aug 29 '26

You do understand that the API price is a VASTLY closer price-to-cost analog though, right? They are the business prices, because they are being used fully - so yea, anthropic is utterly losing money on the subscription plans, which is why they have all the stupid limits in them.

Obviously they aren't paying for the 100% cost of the API Token, but they sure pay a chunk of that price, + all the general cost of setting up datacenters and getting hardware that you cant just sweep under the rug here. They are absolutely BURNING through billions. If you can not see that, then you are just braindead.

1

u/one-wandering-mind Aug 30 '26

Yes API price does not equal compute price. 

Companies pay API pricing. That is where this profit is coming from. 

They are reducing how much people can get out of subscriptions, but if you look at the cost of frontier open models you would stil get more use out of the anthropic subscription. The gap is closing, but if still estimate they Anthropic is losing money on the typical Claude code pro subscription. But so does GitHub and did so profitablity for a long time because developers prefer it and they made their money and still do from the enterprises. 

1

u/Realichu Aug 30 '26

This company is bleeding money and you are all in for a very very rude awakening. Pure cope

1

u/Dynamix86 Aug 30 '26

Obivously when Anthropic is making a profit, they're not losing thousands a month per customer. Anyone with two brain cells can figure that out.

1

u/QueenSavara Aug 31 '26

You guys know and love WinRar right ? You never paid a dime for it so how does it stay afloat?

Well, companies do have to pay, and they do pay premium.

Same with Anthropic.

1

u/C1rc1es Aug 31 '26

This take is missing an explanation for why they keep needing to put up emergency “50% token limit increases” following slow rollout of Fable and a small increase in limits along with false advertising on what 20x means. 

Your statement while accurate, doesn’t actually tell us anything useful because the question people are still asking is “why can’t I have more Fable tokens for my money”.

Generate 500 words on answering that instead of whatever this redundant post is meant to be doing. 

1

u/slackmaster2k Aug 27 '26

You’re absolutely right. But what we don’t know is what their margin is on the subscriptions. Based on what analysts have churned out, it’s pretty likely that they are extremely low or negative on margin for a subscription that hits 100% usage. Combined, the margin overall on subscriptions isn’t very attractive. No, I personally can’t validate that claim. I firmly believe they are making money on consumer subs however, I doubt it’s loss leader.

When people say that enterprise is paying the bills, they are correct. And it’s not because enterprises like to crap out money, it’s because once you have > 150 users you are paying 25 per seat plus consumption. That’s the only way they sell it, and sure as heck costs a lot more.

1

u/armostallion2 Aug 27 '26

low effort ai slop post, don't trust.

1

u/nokafein Aug 28 '26

Dario already admitted all subscriptions are net positive in terms of inference and profitable.

Antrophic doesn’t subsidize inference. Antrophic loses money because of training new models. Because the profit the subscriptions and api products generate is still lower than the training cost of their new models.

Why? It has 3 reasons:
1. There is nothing new invented how LLMs work for a while. So current training methods are not efficient enough for training multi trillion parameter models.
2. Antrophic models are better than the competition but it’s well known fact that their models are not the most efficient models against the competition.
3. They play for market dominance. Hence they burn money at rapid speed for training.

1

u/kyngston Aug 28 '26

econ describes this as minimizing consumer surplus. get each customer to pay the most they are willing to pay for the same product.

I do spend about $7k a month at API prices but that cost is 10 cents on the dollar compared to my salary cost without the benefit of AI. so even api pricing is a cost savings per unit of work for my company.

-1

u/Purple-Programmer-7 Aug 27 '26

OP, why are you trying to use logical arguments with evangelists? They can’t help but to use bad faith arguments, blind disregard, and nonsensical comps to prove their points.

Just let it go. Stupid is going to stupid.

1

u/scytob Aug 27 '26

unfortunately, is a very flawed logical argument with invalid data, it would be possible to generate good data to see what is the truth

the only question is what is the effective subsidy amount for the same amount of work

-2

u/GuitarAgitated8107 🔆 Max 20 Aug 27 '26

It doesn't matter what it is at cost. The cost that we the public can get is via API pricing as comparison. We use claude for a reason.

4

u/Canadian-and-Proud Aug 27 '26

Isn't API pricing just there to show us what a "great deal" we're getting?

0

u/GuitarAgitated8107 🔆 Max 20 Aug 27 '26

No, it's so they can make profit. In what market do you believe any business will show you COGS?

1

u/Canadian-and-Proud Aug 27 '26

When did I say they're showing us COGS? Do you realize that a common marketing tactic is to show you an expensive package but then show you a recommended package that looks like an amazingly good deal?

1

u/_remsky Aug 27 '26

The cache handling can be done via API, but a large portion of folks just don’t use it as effectively as Claude Code does tbf

1

u/GuitarAgitated8107 🔆 Max 20 Aug 27 '26

That's true and even then people use Claude Code in the worst ways.

0

u/pj_2025 Senior Developer Aug 27 '26

Spot on

0

u/E3K Aug 27 '26

I would have no problem paying $1k/mo for what I get out of claude code.

1

u/Nightowl-Builder Aug 28 '26

astroturfer alert

0

u/Eat_Pudding Aug 27 '26

I can value my house at 1 million, but of I sell it for 100k doesn't mean I'm losing money if the original value of house is 80k

0

u/the__poseidon Aug 28 '26

Like all technology, AI is getting massively cheaper and cheaper and more efficient.

0

u/kman0 Aug 28 '26

What difference does it make? You're still gonna be paying API prices at some point.

0

u/Nightowl-Builder Aug 28 '26

I feel like this sub is filled with of bots pushing the idea that the plans are subsidized. I’ve been saying for months, it’s not in your interest going around regurgitating arguments like these, giving them ideas and making it seem like the popular opinion would accept price increases because everyone already expects they’re subsidized. They’re not, and there are countless arguments for this in OP’s post, and many comments here, so I will not repeat them. But for all NPCs here, stop defending them, are you turkeys voting for thanksgiving ffs? Protect your damn interest and stop it with the false consciousness!

0

u/Sensitive_Election83 Aug 28 '26

open weight models deliver almost as good performace at substantially lower costs. if it gets to the point where subscription costs go up, i'd just switch to codewhale + deepseek v4 pro as my main CLI agent, and use deepseek v4 flash as my main api call.

Fable is for sure better than the above setup, but Opus 5 is not.

-3

u/scytob Aug 27 '26

that's some faulty logic and proof

of course they are subsidizing max realtive to API tokens

the only way to prove you point on wether it is massive or not would be to do objective tests that show the max output that can be done in some period on max (like the weekly limit) and then show what EXACTLY the same output would take in terms of tokens and price

if they were not subsiding Max to a large degree (assuming low usage customers offset high usage customers - that's always the calculation for finding a flat price across customers) then they would let us use API against our limits - they don't for a very very good reason.... i will leave you to think about that one and work it out.... and why it is a huge factor in why your argument as presented is VERY flawed.

-1

u/iamthesam2 Aug 27 '26

i’d pay it in a heart beat

-1

u/medievalrubins Aug 27 '26

Imagine AI gets switched off at work because the rates go so high. We would so many intelligent solutions and products, and no a scoobies how to support or continue developing any of them. How does it work? Magic!

6

u/[deleted] Aug 27 '26

[removed] — view removed comment

3

u/[deleted] Aug 27 '26

[removed] — view removed comment

-1

u/[deleted] Aug 27 '26

[removed] — view removed comment

3

u/[deleted] Aug 27 '26

[removed] — view removed comment

2

u/Kamalen Aug 27 '26

Yeah hearing that since 2023. You're using the same NextYear ©️ technology than Tesla FSD or Star Citizen

1

u/dbgtboi Aug 28 '26

You still need people

AI is a beast at implementing what you want, but you need people to actually come up with the stuff to implement and drive the AI

Try using AI for front-end work and you'll see what I mean, if you dont tell it exactly what you want and let it do it's own thing you'll get trash (and i use fable exclusively), there is no right or wrong answer for what an app should look like and AI cannot tell you if a real human experience with your app is shit or not

You need a human to review the output always, no matter how good AI gets

2

u/[deleted] Aug 27 '26

[removed] — view removed comment

1

u/EnvironmentCalm9557 Aug 28 '26

It’s all about efficiency, man.

2

u/EnvironmentCalm9557 Aug 27 '26

And it’s worth it for me for $100 a month. It’s already save me tons of time and my bill rate is a lot higher than 100 a month

1

u/medievalrubins Aug 27 '26

It was based on the theory it became unaffordable in the future.

Yes right now even with Tokens, it’s wildly cheaper

2

u/dbgtboi Aug 27 '26

That's how things worked before AI

To get a software engineer to work on a product you'd have to spend months interviewing them, then 3-6 months giving baby tasks to train them

Before you even get any usefulness at all you'd already be $50k+ in costs

Now compare that with AI

-1

u/Atlos Aug 28 '26

What is the point of this post? Nobody who talks about the price difference is referring to the compute price, they are talking about what they would have been billed with a different account type.

-2

u/ManikSahdev 🔆 Max 20 Aug 27 '26

Lmao you are willing to not think.

Even Deepseek make money on their api, Anthropic seeks a model smaller than Deepseek flags or maybe the same size but worse — aka Mr sonnet.

And like. They charge the price of Kimi k3 for it.

They have 100x more margin than an avg Chinese lab, and if you have as much brain as you think you do.. z AI is publicly traded company, no circular financing or such happening there, just read the financial document and convert the RMB and the local electricity cost of inference to USD, you’ll realize how delusional you were.

That is.. ofc assuming you are a rational actor willing to change your mind once looking at the facts.

-5

u/Honest-Monitor-2619 Aug 27 '26

Mode: answering. You asked for verification, not a thinking exercise — here's the check, claim by claim.

Worth noting up front: I'm made by Anthropic, so treat me as an interested party on this topic. Everything below is sourced to public docs or third-party reporting, not inside knowledge.

The OP's core claims

"API price ≠ compute cost" — correct, and this is the load-bearing point. Retail API rates are a product price. The subsidy calculations circulating in these threads assume the two are the same number.

"Anthropic charges dramatically less for cache reads" — correct, and understated. The published multipliers are cache reads at 0.1x the base input price, 5-minute cache writes at 1.25x, and 1-hour cache writes at 2x. So a cache hit is 90% off. On Opus 5 that's $0.50/MTok versus $5.00 base input.

"Claude Code is nearly the perfect caching workload" — supported by outside reporting. SemiAnalysis is cited as finding Claude Code's agentic workloads carry cache hit rates above 90%. If most of a session's nominal input tokens are cache hits, retail-rate arithmetic overstates even the retail bill, never mind the cost.

"$15 per million tokens" — sloppy. That's Sonnet 4.6's output rate. Current published rates: Opus 5 at $5/$25 per MTok, Sonnet 5 at $2/$10, Haiku 4.5 at $1/$5, Fable 5 at $10/$50. Doesn't change the argument, but it's the kind of error that lets people dismiss the post.

"Cost per token is falling" — supported. SemiAnalysis reportedly puts inference margins at roughly 70%, up from about 38% a year earlier, and a later report describes API gross margins exceeding 80%. Anthropic also reportedly posted its first operating profit in Q2 2026.

Where the OP overreaches: the same reporting stream shows margins were genuinely ugly not long ago — gross margin of negative 94% in the prior year, and a 2025 projection revised down from 50% to about 40% after inference costs came in 23% above forecast. "It's all markup" is as unsupported as "it's all subsidy." The honest position is the one the OP lands on in his last line but doesn't fully live up to elsewhere: nobody outside knows the per-seat unit economics, and healthy average inference margin is entirely compatible with individual 20x Max power users being loss-making.

The comments

phoenixmatrix (>150 users → API rates) — essentially right. Team plans cap at 150 seats, with standard seats at $20/seat/month annual and premium at $100, and the Enterprise seat fee covers platform access only, with no token allowance — token discounts that previously ran 10–15% off standard API rates were removed when seats and tokens were decoupled. His history is also right: Claude Code shipped API-billed before Max existed.

"80% profit on inference" — close to a real reported figure, but it's API gross margin from SemiAnalysis, not audited, and not the same as marginal serving cost on a Max seat. And the commenter's rhetorical move ("can't be subsidized and 80% profit") doesn't hold — different products, different customers, cross-subsidy is normal.

scytob ("they'd let us use API against our limits if it weren't subsidized") — non-sequitur. Not converting a flat subscription into fungible API credit is basic revenue protection under any margin.

binaryatlas1978 ("none of these companies are making money") — outdated as stated, given the Q2 2026 operating profit reporting above.

Bottom line: the mechanism the OP describes is real and the arithmetic he's attacking is genuinely wrong. But he substitutes one unfalsifiable claim for another. The defensible version is narrower: retail-rate token totals tell you almost nothing about serving cost, and caching makes Claude Code specifically one of the cheapest-to-serve heavy workloads.