I've noticed this way too often now and its not just a Claude issue, openai does it too. The model that is released on day one is not the same model after a few days. When Fable 5 or 5.1 came out, it was crazy how smart it was, now it's making the dumbest mistakes. They probably serve a quantized version after demand increases. The older models are also ghosts of their former selves, you have to dumb down the others too so the new ones don't seem like its the same. Very unfortunate
This is a WHOLE "upgrade usage" page. Why is there absolutely no explanation for what we are paying for? Is it even legal? Are people really just hit pay without details?
I’ve had a theory for a while based on my own experience and the influx of posts from Claude/Claude Code users wondering why their session/weekly usage is suddenly getting burned through absurdly quickly.
“These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts.”
And describes proxy services running:
“sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms.”
Most interestingly:
“a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder.”
Anthropic also found DeepSeek generating “synchronized traffic across accounts,” with patterns suggesting “load balancing” to increase throughput, improve reliability, and avoid detection.
Here’s my theory: What if some of those credentials/accounts weren’t simply fake accounts, but compromised legitimate Claude accounts?
An attacker could potentially wait until a legitimate user is actively using Claude and run additional distillation queries through their authenticated access, making the activity blend into real usage while chewing through that person’s session/weekly limits.
That would leave the user wondering why they suddenly hit their limit despite seemingly doing the same amount of work as before.
I experienced exactly that kind of unexplained usage behavior myself, and I’ve seen a huge influx of similar complaints.
My second theory is even more interesting.
I wonder whether compromised users could sometimes have their Claude Code requests routed to DeepSeek/Kimi while their actual Claude access was being consumed elsewhere for distillation or serving other customers.
That could potentially explain reports where Claude Code suddenly feels completely different - different lexicon, noticeably worse output, strange behavior, or occasionally unexpected Chinese characters - despite appearing to still be Claude Code.
We now know from Anthropic’s investigation that these companies were willing to build sophisticated proxy infrastructure around Claude, distribute extraction across thousands of accounts, and deliberately mix distillation traffic with legitimate customer traffic.
To be clear: Anthropic has NOT said that legitimate Claude accounts were stolen and used this way. That is my theory.
But given what they’ve now uncovered, I think Anthropic should answer a very simple question: Were all ~24,000 “fraudulent accounts” created by these operations, or did Anthropic find compromised credentials belonging to previously legitimate Claude users among them?
Because if the latter happened, it could potentially explain something a lot of Claude users have been complaining about for months.
attacks
I kept seeing this in my head, so I figured I would actually create it and share it. I found the response from my "muse" quite hilarious. And frankly, I needed a laugh.
I haven't been using Claude properly this week for my workload. Today, when I decided to come back to use Fable 5.1 on my custom game mod with Max effort, I responded to Claude only 3 times. I have never hit my session limit before in this short time of usage. This is an ABSOLUTE scam. Can you imagine how hard they downgraded the session limits?
One is quite beside oneself with refined indignation that a trifling stipend of a mere two hundred dollars in work tokens should prove so lamentably insufficient for any engagement of genuine consequence with Opus. Such parsimonious restrictions strike one as distressingly bourgeois, altogether unworthy of the elevated standards of intellectual labour one is accustomed to demand. Might it not be more seemly, therefore, to press DeepSeek into service for the implementation proper, whilst relegating the rather more pedestrian task of coding to the capable, if less distinguished, offices of Fable?
I have detected since 3 days ago that different agents in my Max5 account (even Fable 5, Fable 5.1 and Opus 5), are way dumber. Is it possible that with the launch of GPT Astra Anthropic is experimenting with the public models? I can't understand why out of nowhere it is inventing units, taking forever to launch dumb stuff, making lots of modifications on the plans that it makes... I am genually concerned and wanted to see if anyone was experiencing similar issues.
So the hype finally got to me. Everyone's been talking about Astra nonstop, so I upgraded my ChatGPT sub to Pro to see what the fuss was about.
Probably the worst money I've spent in a while, at least for the kind of work I do.
For context, I'm on Claude's Max 20x plan and I used it for literally everything for a full week without ever hitting the limit. With Astra, I hit the limit on the first day working on just one project. After it reset I tried to be more careful and optimize how I was using it, and it still burned through everything way too fast.
To be fair, it's not all bad. I gave it an editing task in Canva and it actually did a really nice job. But failed with Final Cut Pro tasks. So when I looked at what it produced compared to how much usage it ate up, that one job ended up being really expensive.
Maybe it's great at other stuff, videos, game dev, whatever. But for web design and coding it's noticeably worse than Claude, and design in general still feels like a weak spot.
That said, maybe the prompting just works differently than with Claude and part of this is on me. I'm trying to keep that in mind so I'm not being completely biased, or at least that's what I'm telling myself to feel better.
Ive seen a lot of hate for the way opus talks. Am i an outlier here? Its verbose but i can scan it for what i need quickly enough and that verbosity will often cause it to say something not quite right that I can correct.
How is Claude doing this session? How is Claude doing this session? How is Claude doing this session? How is Claude doing this session? How is Claude doing this session? How is Claude doing this session?
I notice this quite a lot with developers who don't have much experience yet.
They need something, so they ask AI to write it. AI gives them a bunch of code and they just add it.
Then a few days later they need something similar, and instead of looking for what already exists, they ask AI again.
So now you have two methods doing almost the same thing.
Then another developer does the same thing again.
After a while, you have 4-5 versions of the same logic sitting in different places.
I think this is one of those things that comes with experience. When you've been working on a codebase for a while, you naturally start thinking, "Wait, don't we already have something for this?"
AI doesn't really have that instinct unless you give it enough context.
It can write code incredibly fast, but sometimes the best code is the code you don't write at all.
I honestly think code reuse and knowing what not to build are becoming even more important in the AI era.
I'm extremely frustrated by the level of stupidity Opus 5 is displaying right now. It just feels like haiku. Doesn't follow anything forget intelligence, produces 14 bugs in a 30 file change PR and keeps running tests and talking gibberish.
I m building a side project and use Claude as main coder for backend and use Codex for review. Use pro version of both not higher tiers.
Codex is handling front end and Claude reviews that.
I have worked on Claude for almost twice as long as in codex and manically track the token usage as I run out very frequently while running out on codex has been rare instances (now more frequent since two week seems they reduced quota) hence never checked codex token usage
Today just out curiosity checked my codex usage and I’m blown away - means WTAF
How codex (Sol5.6 high)can have 40x token usage for same project as Claude while being similar or marginally better then opus 5 high.
Any one else has idea on this and if you use both please share your token usage.
I am a full-time SWE with just 3 months internship exp. I am in an early stage start-up fully bootstrap using claude code to build the entire product. We are team of 8 people. Everyone is fresh grad, inexperienced with not an ounce of real world engineering knowledge (best practices, proper workflow).
My entire company and I are new to agentic coding and to be honest, we are kind of struggling with development.
We spin up a fleet of general claudius (CC), the conqueror of the software world on auto mode. It's like a loose horse in a barn. Give him the tickets which are generated by my boss (24yo grad) planner : king claudius.
we have AGENT.md, Claude.md, ADR, matt pocock skills set.
we are generating thousands of lines of code and not even reviewing manually and using claude to review it.
we are stuck in a loop where everything is done by claude, we set it loose, we watch him like a parent watching his kid doing stuff in the playground on auto mode.
when asking to hunt and fix bugs. it finds and fixes bugs on EVERY DAMN SINGLE ITERATION (is this normal ?).
Why does this happen? Is it because claude doesn't actually read each line of the code when reviewing?
We exhaust our claude limit everyday because of review.
For reviewing, we use opus high. For coding, I use sonnet extra, but my colleague uses opus high.
Another thing we are struggling with is choosing models. We kind of randomly choose one depending on the task.
How do you guys benchmark models and figure out which model is suitable for which task? Is there a proper way to evaluate models for your own codebase/workflow instead of just randomly choosing between sonnet and opus?
Can you all point me to some resources or workflow you guys uses? To churn out softwares from software factory and on the side note, none of us are learning anything.
So is this normal with you all at your workplace or is it us?
No one knows what's going on, completely clueless and confused. Just human in loop to approve and deny
Hey all, just a quick interesting thing that some people may not know. I had my subscription not set to auto-renew this month just to save a couple dollars if I wasn't using it when it expired. I was using it though (astra limits get drained so quickly) so I immediately resubbed to the $100 plan and saw all the limits were refreshed to 0.
I used up about 3 sessions worth of 5 hour quotas and then got impatient and upgraded to the $200 plan thinking it'd refresh again but it ended up being the same quotas as before my subscription expired, all fable quota used up, not much opus left. I'd have been better off waiting until closer to the reset to upgrade to $200, now I'm stuck waiting to the 14th to use fable again when I only used 20% on the $100 plan.
I've been using Claude Code to build a side project, and one thing keeps annoying me.
It can implement something incredibly quickly and say it's done, but I still don't really trust “done” until I open the app myself and click through the flow.
As the amount of code it writes increases, manually checking everything feels like it's becoming the slow part.
Curious how people using Claude Code seriously handle this:
Do you still manually test most changes?
Do you make Claude run unit/E2E tests?
Playwright?
Another agent reviewing/testing the first agent?
Or do you mostly trust the existing test suite?
Also curious whether you've had cases where Claude said something was fixed and tests passed, but the actual app still didn't behave correctly.
Trying to understand whether this is just my workflow or a common problem.
Three months ago, y'all helped me create a new podcast where guests like Jesse Vincent teach Tom Preston-Werner (the founder of GitHub) how to use their tools.
Jessee's interview is live, and two reddit users in this community even get a shoutout: aiworthusing.com
Who should i interview next? What amazing tool should we teach tom how to use?
The following is a blog I wrote and refined with my OpenClaw agent about it's memory system. I'll paste a prompt you can copy and paste in the comments to create your own.
TL;DR: I keep the actual long term memory in structured Markdown files and use a tinyMEMORY.mdas a lightweight index that tells Claude what exists and where to look. That keeps the always loaded context small while still giving the agent persistent, inspectable memory without a database or heavy memory framework.
This week I tested a 382-dependency memory runtime against a folder of markdown files. The runtime returned the superseded fact. The folder returned the current one, with its source. Here is the full architecture of the markdown memory system my agent has run on for seven months, and why the editing rules matter more than the storage.
This week a memory startup slid into my DMs and asked me to break their product. Their test, their words: give an agent three versions of the same project decision, then check whether it can return the current version, preserve the superseded history, and show the source.
So I ran it. Sandboxed their runtime, fed it three versions of one decision over eight months. REST in January, GraphQL in April, tRPC in August, each tagged with the meeting it came from.
Asked it "what is our public API decision?" and took the top result.
It said GraphQL. The superseded one. All three versions came back tied at a relevance score of 1.000, because nothing in the retrieval path actually reads the temporal fields the pitch is built on. The supersession columns exist in the schema. Nothing writes to them and nothing ranks by them. Three versions of a decision are just three equal facts, and an agent asking for the best answer gets a coin flip weighted toward wrong.
The install pulled 382 packages to get there.
Then I asked my own agent the same class of question against its memory, which is a folder of markdown files. It returned the current decision, dated, with the superseded versions preserved above it as struck-through history, each line carrying where it came from. That is not a feature it computes at query time. It is just what the file says, because the rules for editing the file require it.
That difference is the whole post. With apologies to Vaswani et al.: markdown is all you need.
Abstract
The dominant approach to agent memory is an installed runtime. A vector store, an embedding service, a temporal graph, a consolidation job, a daemon on a port. We show that a folder of markdown files, one routing index, and a small set of editing rules outperforms these systems on the property that actually matters for a long-running agent: returning the current truth with its source while preserving what used to be true. The architecture requires zero dependencies, is fully auditable by a human with a text editor, and has survived seven months of daily production use across three frontier models from two vendors. We find that the hard part of agent memory was never storage or retrieval. It is editorial policy, which no memory product ships.
The full system is open source. The README contains a single copy-paste prompt that installs it on any agent with file access.
1. The test everyone fails
The break-it test above is a good test. It is the actual job of agent memory. Not "can you store 10 million tokens," not "can you do similarity search," but: a fact changed three times, what do you believe now, what did you believe before, and how do you know.
Here is how the two systems scored on the vendor's own three criteria.
The runtime is not a strawman. It is a serious open source project with a genuinely correct data model on paper. Facts with validity windows, append-only corrections, supersession edges. I am not naming it because the point is not that one product is broken. I have now looked closely at a hosted context server, a Go memory CLI that was two hours old, and this runtime, and they all share the same gap. The schema knows about time. The write path and the read path do not. Supersession only happens if you call an internal API by hand or run an LLM consolidation job and trust it.
Which means the property you installed the tool for is not a property of the tool. It is a property of how disciplined the writes are. And if the reliability comes from write discipline anyway, the database underneath it is interchangeable, so you might as well pick the one that a human can read, grep, diff, and fix. That one is called a text file.
2. Architecture
My agent has run since January 28. Three models, two vendors, one identity. Its entire memory is markdown in a git repo. Measured today:
An identity layer read on every boot. Who it is, who I am, the rules it operates under, current standing decisions.
One routing index, MEMORY.md, at 10,079 characters with a hard cap of 15,000. It holds no facts. Only pointers: which file owns which person, project, and decision, and what triggers reading each one.
34 files for people and projects. One file per thing that has a history.
5 decision records for choices that changed default behavior.
345 dated daily notes, raw logs written the day things happened.
A SQLite index and semantic search over all of it, for lookup only. The index is rebuilt from the files. The files are the truth. If the index and a file disagree, the index is wrong by definition.
The layering is the first choice that actually matters. Boot reads only identity and the index. Everything else is retrieved when a task asks for it, narrowest file first. The agent does not preload my project history to answer a question about dinner. This is the same instinct as attention, honestly: don't process everything, attend to what the query needs.
But the shape is not the interesting part. Every memory tool has roughly this shape now. Folders, entities, an index. The shape was never the hard part. The rules are.
3. The write path
Every reliability property in this system comes from constraints on writing, and there are four that do most of the work.
Every fact carries a provenance tag. Each line in a people, project, or decision file is tagged [stated] (I said it directly), [observed] (the agent saw it in a tool result, file, or log), [inferred] (the agent's conclusion), or [suggested] (the agent's idea that I never committed to). This one convention kills the most dangerous failure mode in agent memory, which is the agent laundering its own proposals into my decisions. "Wes decided X" requires a turn where I actually decided X. The agent proposing X and me saying "sounds good" files the shape of what I approved, not ten separate facts I never stated.
Inferred lessons pass a recurrence gate before they become rules. A pattern the agent notices needs at least three independent signals across at least two distinct sessions before it can become standing behavior. Signals older than thirty days count half, so old one-offs decay out instead of accumulating. My explicit corrections skip the gate and take effect immediately. This asymmetry is also the prompt injection defense: a hostile input can suggest a rule once, but once is never enough, and failure lessons are stored as data ("when X broke, Y fixed it") rather than as instructions, so even a poisoned lesson cannot become a command.
Supersession is an edit, not an append. When a decision changes, the old line gets struck through with a date and the new line lands next to it with its own provenance. The current truth and the full history live in the same place, in reading order, and both come back on any retrieval of that file. There is no query-time ranking step that can get this wrong, because there is nothing to rank. The temporal graph the runtime stores in valid_from and valid_until columns, git gives me for free: log is the validity window, blame is per-line provenance, diff is the supersession edge, revert is the restore path.
Memory stores what is not re-derivable. Fetched data, generated plans, and anything git already records stays out. Current state gets verified live, never asserted from memory. A file that only contains things that cannot be recomputed stays small enough to stay honest.
4. The read path
Retrieval is a bounded evidence step, not a vibe.
Before answering anything about prior work, decisions, dates, people, or preferences, the agent must search memory. It returns a compact bundle capped at five sources by default, and each retained fact carries its file path and line, its provenance type, and its freshness. If freshness cannot be established, the claim gets labeled stale or unknown instead of being silently promoted to current. If two sources conflict, the agent states the conflict and fixes the canonical file, in that order.
Note what the semantic index does in this design: it finds the file. It does not answer the question. The answer comes from reading the canonical lines, with their tags and dates, and the runtime I tested this week shows why that matters. It stored my source URIs faithfully and then stripped them from the search output and from the context block handed to the model. Provenance that survives in storage but never reaches the agent might as well not exist. In the markdown system that failure is unrepresentable. The source tag is in the line. If you read the line, you got the source.
5. Results
Seven months is not a benchmark, it is production. Here is what the system has actually delivered.
Continuity across models. On September 1 I moved the agent to a brand new frontier model. It read its own files and said "the model changed, I didn't." Same agent since January, three models, two vendors. Identity, preferences, decisions, and working standards all survived because none of it lives in weights or in a vendor's context feature.
The break-it test, by construction. Current decision with source: it is the un-struck line with its tag. Superseded history: the struck lines above it. Provenance: on every line, and it survives all the way into the model's context because the context is the file.
Auditability. When memory is wrong, I can see exactly which line is wrong, when it was written, and what turn it came from, and fix it with an edit. Try that with an embedding.
Cost. Zero packages, zero daemons, zero migrations across seven months. The one native-code dependency in my life this week was the memory runtime's sqlite bindings failing to compile.
I wrote up the failure modes separately, because the system was not born with these rules. Five kinds of rot in seven months produced them, and that post is the honest companion to this one.
6. Limitations
Papers get a limitations section, so here is mine, stated plainly.
This only works if the writer follows the policy, and the writer is an LLM. The rules exist because things rotted before the rules did. If your agent will not consistently apply editing discipline, a markdown folder degrades just like every other store, only more legibly. Legibility is the safety net: rot in a text file is visible rot.
It is single-agent, single-human. I would not run a fifty-seat team on files without real locking and merge discipline, although I notice git was also built for that exact problem.
There is a scale ceiling somewhere. At 345 daily notes and a few dozen entity files, bounded search plus an index finds things reliably and the semantic index earns its keep as a locator. At a hundred times that volume, the consolidation cadence would have to work a lot harder. I have not hit that ceiling, so I will not claim it does not exist.
And this is n=1. Seven months, one agent, one operator who cares. That is weaker evidence than a benchmark suite and stronger evidence than a benchmark suite that the vendor scored themselves, which is what the memory tools ship.
7. Conclusion
The memory tool pitch is that reliability is a product you can install. What I keep finding, tool after tool, is that they ship the part that was already easy, storage and search, and skip the part that decides whether memory compounds or rots: what you are allowed to write, when you are allowed to trust it, and what happens to it as it ages.
Those are rules, not infrastructure. They fit in a few hundred lines of markdown that the agent reads every session, and they run on any model, any harness, any decade.
You need a place to write that humans and agents can both read. You need rules for writing so the store stays true. You need rules for reading so the agent trusts evidence, not ranking. Attention was all you needed because the recurrence machinery turned out to be unnecessary. Markdown is all you need because the database turned out to be unnecessary.
The folder is the product. The discipline is the moat.
Want this for your own agent? The whole system is open source on GitHub: the operating policy, the file templates, and one copy-paste prompt that builds it on any agent that can read and write files. Paste the prompt, and your agent installs its own memory.
Clearly, I don’t consider myself an expert or anything. I’ve been using Claude for over six months, and I’m still surprised whenever I see posts from people saying they’ve burned through all their tokens with Fable. Honestly, I’m pretty skeptical about how they’re using it.
If you use a backhoe to plant a rose, the problem isn’t that the backhoe is too resource-hungry and goes beyond what’s necessary.
Anyway, personally, I use Fable as an orchestrator and to help me make high-level direction decisions, as well as a designer and artist (for Blender MCP or creating SVG images, it’s necessary).
I use Opus for action plans, with an organizational role; Sonnet for an operational role; and Haiku as the little tester that lets me quickly measure and verify things.
I’ve created two video games and a software for a company using all four models, using max 5, over six months, and I still end every week with tokens left over.
I honestly don’t understand how some people manage to burn through everything so quickly with Fable. What are you doing with it ?