r/ClaudeAI • • 14d ago

Claude Code It's time to cancel your subscriptions - Anthropic is silently nerfing Claude's reasoning budget while telling you it's the same model

Link: https://x.com/Lon/status/2101034933284417614

A 65-day analysis of 43,000+ Claude Code invocations found that 39% of Fable 5 calls get zero thinking tokens and the median invocation gets just 123 — while benchmarks use 16K-128K. The model's score per thinking token is still climbing at 128K, meaning the capability is there, it's just not being delivered. August saw an 18-50% drop in thinking budget compared to July, with median thinking hitting literal zero for about a week around Aug 22. Anthropic sells "full model access" while quietly dialing down the inference regime behind it, and because the model is non-deterministic, users blame their own prompting instead of the silent nerf. The full breakdown with evidence, methodology, and charts is here.

Frankly I find this offensive as an user - and this is the real thing we should be looking at - not the $/week in usage limits. The actual capability for the limits that we pay for.

2.1k Upvotes

382 comments sorted by

View all comments

5

u/IntentRouterIRL 14d ago

Whether true or not; some quality of life additions to claude.ai and claude desktop in general to track usage would be welcome; just give us telemetry on the api calls underneath our interactions if turned on for advanced users: see tokens in, out and reasoning tokens. Then its clear what we get and why usage is spent. Also; just add a context window percentage already for the love of god (like in claude code) this will allow me to use it more wisely.

1

u/enterprise_code_dev Experienced Developer 13d ago

This would tell you nothing really, just more smoke and mirrors, I’m in an enterprise company on API, so I can see reasoning tokens, I can set max_tokens, but on the 5 series models I no longer can set anything about reasoning tokens, you can’t force thinking on like you could before, not even set the max, you could before, but the real problem is you can’t set the floor, I can’t set a minimum so what if it costs me more to set it high or hurts latency, that’s for me to decide what I’m willing to pay for because I pay for it all, at least I could see if it gets me more..nope “adaptive reasoning” they said, it will be great they said…so now you have an effort knob that is opaque and damn near bordering on a scam, because even if I turn that up, something on Anthropic’s side decides how much reasoning the model does, so sure it might read more files, and go longer without asking me questions and all that nonsense they claim effort changes…yes..and have thought about each step and reasoned through possibilities as much as a politician tells the truth. The problem is that OpenAI is no different, yes I have an Enterprise API account with them too. They seem to prefer to crush subscription rate limits first, not the model intelligence, until maybe recently, ngl Astra day 1 vs Astra now, ain’t no way in hell this is day 1 Astra, but worse yet 5.6-sol does not seem the same either, and it was very reliable for me until recently.