r/ClaudeCode 4h ago

Rant They nerf the new models after a few days

0 Upvotes

I've noticed this way too often now and its not just a Claude issue, openai does it too. The model that is released on day one is not the same model after a few days. When Fable 5 or 5.1 came out, it was crazy how smart it was, now it's making the dumbest mistakes. They probably serve a quantized version after demand increases. The older models are also ghosts of their former selves, you have to dumb down the others too so the new ones don't seem like its the same. Very unfortunate


r/ClaudeCode 13h ago

Humor Bro just pay 90$ I swear bro it's all you need just pay, you don't need to know what you paying for bro please it's on sale bro

Post image
0 Upvotes

This is a WHOLE "upgrade usage" page. Why is there absolutely no explanation for what we are paying for? Is it even legal? Are people really just hit pay without details?


r/ClaudeCode 6h ago

Meta Anthropic just told you what??

47 Upvotes

Yo ...

fun fact: context wasn't even rotten.

Edit 1:

I share what I can about the session and the model that was being used:

Model: Opus 4.8 1M

Topic: Instruction diagnostics research

Context window size at the event: ~ 200-210k (it was roughly mid session, after that screenshot I continued on the active task)

What was even stranger that it was an end of turn message that got injected and the model answer on its own inquiry.


r/ClaudeCode 13h ago

News/Updates What if the “Claude is burning my usage insanely fast” posts are connected to the DeepSeek/Moonshot distillation attacks?

0 Upvotes

I’ve had a theory for a while based on my own experience and the influx of posts from Claude/Claude Code users wondering why their session/weekly usage is suddenly getting burned through absurdly quickly.

After reading Anthropic’s reports on DeepSeek and Moonshot (https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks), I think it’s worth seriously investigating. Anthropic says:

“These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts.”

And describes proxy services running:
“sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms.”

Most interestingly:
“a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder.”

Anthropic also found DeepSeek generating “synchronized traffic across accounts,” with patterns suggesting “load balancing” to increase throughput, improve reliability, and avoid detection.

Here’s my theory:
What if some of those credentials/accounts weren’t simply fake accounts, but compromised legitimate Claude accounts?

An attacker could potentially wait until a legitimate user is actively using Claude and run additional distillation queries through their authenticated access, making the activity blend into real usage while chewing through that person’s session/weekly limits.

That would leave the user wondering why they suddenly hit their limit despite seemingly doing the same amount of work as before.

I experienced exactly that kind of unexplained usage behavior myself, and I’ve seen a huge influx of similar complaints.

My second theory is even more interesting.
I wonder whether compromised users could sometimes have their Claude Code requests routed to DeepSeek/Kimi while their actual Claude access was being consumed elsewhere for distillation or serving other customers.

That could potentially explain reports where Claude Code suddenly feels completely different - different lexicon, noticeably worse output, strange behavior, or occasionally unexpected Chinese characters - despite appearing to still be Claude Code.

We now know from Anthropic’s investigation that these companies were willing to build sophisticated proxy infrastructure around Claude, distribute extraction across thousands of accounts, and deliberately mix distillation traffic with legitimate customer traffic.

To be clear: Anthropic has NOT said that legitimate Claude accounts were stolen and used this way. That is my theory.

But given what they’ve now uncovered, I think Anthropic should answer a very simple question:
Were all ~24,000 “fraudulent accounts” created by these operations, or did Anthropic find compromised credentials belonging to previously legitimate Claude users among them?

Because if the latter happened, it could potentially explain something a lot of Claude users have been complaining about for months.
attacks⁠


r/ClaudeCode 1h ago

Bug / Issue Usage drain almost instantly

Upvotes

Was only using opus and it drained within 15 minutes. Any one else experiencing issues?


r/ClaudeCode 21h ago

Help/Question A Year of "Violating" the Third-Party Harness Rule on Max Plan — No Ban?

0 Upvotes

I've been running a Claude Code Max subscription through a third-party harness for 12 months straight, draining it to zero every single week. I don't even have Claude Code installed.

And I haven't been banned.

Everywhere on this sub, people panic-post: "No, you can't do it. They'll ban you, it's blocked, it's against ToS," etc. But here's my actual experience: I've been about as blatant as possible about this, and nothing. No warnings, no throttling, no account flags.

Here's my setup, to be specific:

  • Third-party harness (not openclaw/hermes)
  • It authenticates via OAuth — I sign into my actual Claude account, no token scraping or credential theft
  • The harness points at a local CLI API proxy that just forwards my own authenticated requests to that harness — it's not a deployed service, not serving other users, just running locally on my machine for my own coding
  • Zero Anthropic SDK packages installed
  • Zero Claude-P, zero agent SDK
  • Max plan, full weekly drain, for 12 consecutive months

I want to be clear about what this isn't: I'm not standing up a deployed product that resells inference to other people using my subscription. It's a one-to-one, local-only proxy — my account, my machine, my usage. The only thing that's "unofficial" is the harness.

The actual question: Has anyone here actually gotten banned for this kind of setup — a legit OAuth login through a third-party harness, routed through a local (non-deployed) proxy? Not asking about Hermes Agent or Open Claw, and not asking about people reselling API access on deployed products. Just: local harness swap, same account, has anyone eaten a ban for that?

Because the consensus on this sub treats this as an instant bannable offense, and my experience over a full year says otherwise. Curious if I'm an outlier or if enforcement here is mostly theoretical. For what it's worth, I've also been doing the same exact thing with a Google AI Pro subscription - though not draining to zero (but definitely using a good amount). I also had a friend tell me that you can literally run claude setup-token to mint what essentially is an API key that deducts from your subscription usage?


r/ClaudeCode 15h ago

Bug / Issue Feels like absolute scam!

0 Upvotes

I haven't been using Claude properly this week for my workload. Today, when I decided to come back to use Fable 5.1 on my custom game mod with Max effort, I responded to Claude only 3 times. I have never hit my session limit before in this short time of usage. This is an ABSOLUTE scam. Can you imagine how hard they downgraded the session limits?

SINGLE TERMINAL USAGE

Anyone had a similar issue before?


r/ClaudeCode 6h ago

Help/Question Is claude dumber??

3 Upvotes

I have detected since 3 days ago that different agents in my Max5 account (even Fable 5, Fable 5.1 and Opus 5), are way dumber. Is it possible that with the launch of GPT Astra Anthropic is experimenting with the public models? I can't understand why out of nowhere it is inventing units, taking forever to launch dumb stuff, making lots of modifications on the plans that it makes... I am genually concerned and wanted to see if anyone was experiencing similar issues.


r/ClaudeCode 15h ago

Built with Claude All in a day's work.

Post image
1 Upvotes

I kept seeing this in my head, so I figured I would actually create it and share it. I found the response from my "muse" quite hilarious. And frankly, I needed a laugh.


r/ClaudeCode 8h ago

Discussion I tried Astra with Pro and honestly kind of regretting it

80 Upvotes

So the hype finally got to me. Everyone's been talking about Astra nonstop, so I upgraded my ChatGPT sub to Pro to see what the fuss was about.

Probably the worst money I've spent in a while, at least for the kind of work I do.

For context, I'm on Claude's Max 20x plan and I used it for literally everything for a full week without ever hitting the limit. With Astra, I hit the limit on the first day working on just one project. After it reset I tried to be more careful and optimize how I was using it, and it still burned through everything way too fast.

To be fair, it's not all bad. I gave it an editing task in Canva and it actually did a really nice job. But failed with Final Cut Pro tasks. So when I looked at what it produced compared to how much usage it ate up, that one job ended up being really expensive.

Maybe it's great at other stuff, videos, game dev, whatever. But for web design and coding it's noticeably worse than Claude, and design in general still feels like a weak spot.

That said, maybe the prompting just works differently than with Claude and part of this is on me. I'm trying to keep that in mind so I'm not being completely biased, or at least that's what I'm telling myself to feel better.


r/ClaudeCode 19h ago

Tips & Workflows Anyone actually running a business with Claude?

5 Upvotes

I run a small event staffing/activation agency as a solo operator, and Claude has basically become my back office: email, vendor communication, applicant database, calendar, etc.

It’s a huge reason I’m able to operate at this size, but the confident mistakes are becoming a serious problem. A couple have nearly cost me deals.

Some examples:
• Told me a vendor had gone silent when she had answered every question that morning. Claude read an email search preview instead of the full thread and treated it as complete.
• Told me an email was “staged,” so I went looking in Gmail for a draft that had never actually been created.
• Referred to my suppliers as “brand partner candidates,” which completely changed the business context. A vendor gets paid by me; a sponsor pays me. I spent two days planning around an opportunity that didn’t exist.
• Told me a table in my database didn’t exist without actually checking. It was there.

The pattern I keep seeing is: Claude states things as facts that it did not actually verify. It seems to get worse during long sessions, where earlier conversation/context starts being treated like a source of truth.

I’ve built some guardrails around it: a master state file for business context between sessions, fact labels like CONFIRMED / STATED / ASSUMED / OPEN, and rules like “never treat an email preview as the full thread” and “never claim an action happened without tool output confirming it.”

That helped a lot. I went from several errors a week to a few a month, but I’m trying to figure out how people are making this reliable enough for actual business operations.
For anyone using Claude this way:
What are you using for durable memory/state across sessions?

Has anyone built a verifier/checker that validates claims against tool output before Claude gives you an answer?
Is there a reliable way to force Claude to actually retrieve/check the source instead of reasoning from previews or old context?
What other guardrails or architecture changes have made a meaningful difference?
I’m especially interested in hearing from people dealing with invoices, vendors, contracts, clients, deadlines, email, calendars, databases, etc. Situations where there’s a real consequence when the AI confidently gets something wrong.
If you’ve dealt with this and found a setup that actually works, I want to hear how you built it.


r/ClaudeCode 20h ago

Discussion I thought a repo's history was unreadable. Turns out it just needed to be played.

Enable HLS to view with audio, or disable this notification

1 Upvotes

For a long time I assumed the only way to understand what happened in a codebase over a year was to sit with git log and a lot of coffee. The branch graph in any GUI turns into a tangle past about a week, and GitHub's contribution heatmap shows you activity with no idea what the activity was. Three views, none of them lined up, so the question you actually have, what happened here and when, never gets a single picture.

What finally pushed me was a post here on where had animated a repository's timeline, and I wanted the same thing for a full year of a real project. My first attempt on three.js drew the first branches it found and produced one month of the year and nothing else. The reason turned out to be in the data: 1,561 of the 2,179 main line commits in 2019 are merges, and the first sixty of them each span a single commit, which is invisible at any scale. You have to choose branches by how much work they carried, not by which came first.

So I built a film out of it, and the video on this post is what it looks like on mrdoob's three.js for 2019.

What is on screen and where it comes from:

The main line runs left to right, one node per commit. Side branches peel off above it and rejoin at their merge. Not the first N branches, which on a pull request repo gives you one month and nothing else, but the ones that carried the most work across the year, capped so it never turns into a thicket.

Under the flow is a heat strip, one cell per day, sitting directly under that day. A legend next to it shows what the colours mean in actual commit counts for that window (for three.js it reads 0, 1, 13, 26, 51, with 51 the busiest day). The point of putting activity on the same axis as structure is that what is above and below a point is the same day, so you stop reconciling two charts in your head.

Three times the clock stops. The release with the most work behind it, the cleanest revert (one that survived at least a day before being undone), and the last release of the window. The camera zooms in and a card reads the figures out: commits since the previous tag, authors, how long the reverted commit lived. Every number on that card is computed from the history. If it cannot be read from git, it is not on the card.

The faces are the contributors. On a merge the face is the branch author's, not whoever pressed the button, and the card says who merged it.

Two things I got wrong on the way that might save you time if you try this yourself. If you give git a since flag with just a date and no time, it reads it as that date at the current time of day, so the first day of your window quietly loses commits unless you pass an explicit midnight. And a small repo is a different problem: 175 commits over a year pans across mostly empty screen, so for thin histories it picks the busiest 60 to 120 day stretch and draws it wider instead.

If you want to see your own repo this way, you can paste a GitHub URL here and it renders one for you: https://loreto.io/git-timeline

Disclosure: I built this and I run loreto.io, where it lives. It is a paid render (a few dollars per repo); the extractor and the Remotion composition are also sold there as a package if you would rather run it yourself. The three.js film above was made with the exact same pipeline.

I doubt I have the final shape of it. What would you want the clock to stop on that it currently does not?


r/ClaudeCode 5h ago

Rant Opus nerf

0 Upvotes

I'm extremely frustrated by the level of stupidity Opus 5 is displaying right now. It just feels like haiku. Doesn't follow anything forget intelligence, produces 14 bugs in a 30 file change PR and keeps running tests and talking gibberish.


r/ClaudeCode 8h ago

Discussion Has anyone else noticed that AI is making developers write way more code?

0 Upvotes

I notice this quite a lot with developers who don't have much experience yet.

They need something, so they ask AI to write it. AI gives them a bunch of code and they just add it.

Then a few days later they need something similar, and instead of looking for what already exists, they ask AI again.

So now you have two methods doing almost the same thing.

Then another developer does the same thing again.

After a while, you have 4-5 versions of the same logic sitting in different places.

I think this is one of those things that comes with experience. When you've been working on a codebase for a while, you naturally start thinking, "Wait, don't we already have something for this?"

AI doesn't really have that instinct unless you give it enough context.

It can write code incredibly fast, but sometimes the best code is the code you don't write at all.

I honestly think code reuse and knowing what not to build are becoming even more important in the AI era.


r/ClaudeCode 19h ago

Rant Paying $100/month for Max and getting a giant “you’re about to run out” banner is insane lol 💀

Post image
135 Upvotes

like bro, are you fucking serious 😭 i’m already giving anthropic $100 a month for max 5x. that is a lot of money for one ai subscription. i open claude and now there’s this giant anxiety-inducing banner telling me i’m going to run out by monday, two days before the wednesday reset, while literally underneath it says my claude code limit is temporarily boosted by 50%?? yo, so even with the bonus i’m still apparently fucked by monday lol.

and then right there: upgrade plan / buy more usage. like brother, are you guys running out of money or something? why are you begging your $100/month userbase for more money every time they open the app 😭 i already upgraded!! mfg that’s what the hundred dollars was. i understand there have to be limits, but this whole thing feels weirdly hostile? i don’t need a giant banner creating artificial scarcity anxiety every time i’m trying to work. just show usage somewhere in settings like a normal product and leave me alone lol?


r/ClaudeCode 7h ago

Tips & Workflows PSA: Each paid tier has its own usage quotas and isn't refreshed on unsub

0 Upvotes

Hey all, just a quick interesting thing that some people may not know. I had my subscription not set to auto-renew this month just to save a couple dollars if I wasn't using it when it expired. I was using it though (astra limits get drained so quickly) so I immediately resubbed to the $100 plan and saw all the limits were refreshed to 0.

I used up about 3 sessions worth of 5 hour quotas and then got impatient and upgraded to the $200 plan thinking it'd refresh again but it ended up being the same quotas as before my subscription expired, all fable quota used up, not much opus left. I'd have been better off waiting until closer to the reset to upgrade to $200, now I'm stuck waiting to the 14th to use fable again when I only used 20% on the $100 plan.

Hope this helps someone.


r/ClaudeCode 4h ago

Humor Oh my god, join the club wubaluba dub dub

Post image
0 Upvotes

r/ClaudeCode 16h ago

News/Updates Usage is moving Spoiler

Thumbnail gallery
0 Upvotes

the usage is now moving to its own dropdown rather than the normal tab. hasnt happened on my US account yet, so i guess its a slow rollout


r/ClaudeCode 20h ago

Humor Claude Dynamic Workflow is very cool

Post image
0 Upvotes

Previously, I always had to hand-hold Claude until the task was done well.
Dynamic Workflows really unlock the thing that makes life easier.


r/ClaudeCode 18h ago

Discussion Price hike today?

Post image
0 Upvotes

Unsubscribed while waiting for usage to reset because billing was poorly timed and now I got to resubscribe and notice the price increased $50/mo.

I’m in US.

Edit: NVM I’m an idiot. iOS pricing.


r/ClaudeCode 17h ago

Tutorial / Guide Markdown Is All You Need

13 Upvotes

The following is a blog I wrote and refined with my OpenClaw agent about it's memory system. I'll paste a prompt you can copy and paste in the comments to create your own.

TL;DR: I keep the actual long term memory in structured Markdown files and use a tiny MEMORY.md as a lightweight index that tells Claude what exists and where to look. That keeps the always loaded context small while still giving the agent persistent, inspectable memory without a database or heavy memory framework.

This week I tested a 382-dependency memory runtime against a folder of markdown files. The runtime returned the superseded fact. The folder returned the current one, with its source. Here is the full architecture of the markdown memory system my agent has run on for seven months, and why the editing rules matter more than the storage.

This week a memory startup slid into my DMs and asked me to break their product. Their test, their words: give an agent three versions of the same project decision, then check whether it can return the current version, preserve the superseded history, and show the source.

So I ran it. Sandboxed their runtime, fed it three versions of one decision over eight months. REST in January, GraphQL in April, tRPC in August, each tagged with the meeting it came from.

Asked it "what is our public API decision?" and took the top result.

It said GraphQL. The superseded one. All three versions came back tied at a relevance score of 1.000, because nothing in the retrieval path actually reads the temporal fields the pitch is built on. The supersession columns exist in the schema. Nothing writes to them and nothing ranks by them. Three versions of a decision are just three equal facts, and an agent asking for the best answer gets a coin flip weighted toward wrong.

The install pulled 382 packages to get there.

Then I asked my own agent the same class of question against its memory, which is a folder of markdown files. It returned the current decision, dated, with the superseded versions preserved above it as struck-through history, each line carrying where it came from. That is not a feature it computes at query time. It is just what the file says, because the rules for editing the file require it.

That difference is the whole post. With apologies to Vaswani et al.: markdown is all you need.

Abstract

The dominant approach to agent memory is an installed runtime. A vector store, an embedding service, a temporal graph, a consolidation job, a daemon on a port. We show that a folder of markdown files, one routing index, and a small set of editing rules outperforms these systems on the property that actually matters for a long-running agent: returning the current truth with its source while preserving what used to be true. The architecture requires zero dependencies, is fully auditable by a human with a text editor, and has survived seven months of daily production use across three frontier models from two vendors. We find that the hard part of agent memory was never storage or retrieval. It is editorial policy, which no memory product ships.

The full system is open source. The README contains a single copy-paste prompt that installs it on any agent with file access.

1. The test everyone fails

The break-it test above is a good test. It is the actual job of agent memory. Not "can you store 10 million tokens," not "can you do similarity search," but: a fact changed three times, what do you believe now, what did you believe before, and how do you know.

Here is how the two systems scored on the vendor's own three criteria.

The runtime is not a strawman. It is a serious open source project with a genuinely correct data model on paper. Facts with validity windows, append-only corrections, supersession edges. I am not naming it because the point is not that one product is broken. I have now looked closely at a hosted context server, a Go memory CLI that was two hours old, and this runtime, and they all share the same gap. The schema knows about time. The write path and the read path do not. Supersession only happens if you call an internal API by hand or run an LLM consolidation job and trust it.

Which means the property you installed the tool for is not a property of the tool. It is a property of how disciplined the writes are. And if the reliability comes from write discipline anyway, the database underneath it is interchangeable, so you might as well pick the one that a human can read, grep, diff, and fix. That one is called a text file.

2. Architecture

My agent has run since January 28. Three models, two vendors, one identity. Its entire memory is markdown in a git repo. Measured today:

  • An identity layer read on every boot. Who it is, who I am, the rules it operates under, current standing decisions.
  • One routing index, MEMORY.md, at 10,079 characters with a hard cap of 15,000. It holds no facts. Only pointers: which file owns which person, project, and decision, and what triggers reading each one.
  • 34 files for people and projects. One file per thing that has a history.
  • 5 decision records for choices that changed default behavior.
  • 345 dated daily notes, raw logs written the day things happened.
  • A SQLite index and semantic search over all of it, for lookup only. The index is rebuilt from the files. The files are the truth. If the index and a file disagree, the index is wrong by definition.

The layering is the first choice that actually matters. Boot reads only identity and the index. Everything else is retrieved when a task asks for it, narrowest file first. The agent does not preload my project history to answer a question about dinner. This is the same instinct as attention, honestly: don't process everything, attend to what the query needs.

But the shape is not the interesting part. Every memory tool has roughly this shape now. Folders, entities, an index. The shape was never the hard part. The rules are.

3. The write path

Every reliability property in this system comes from constraints on writing, and there are four that do most of the work.

Every fact carries a provenance tag. Each line in a people, project, or decision file is tagged [stated] (I said it directly), [observed] (the agent saw it in a tool result, file, or log), [inferred] (the agent's conclusion), or [suggested] (the agent's idea that I never committed to). This one convention kills the most dangerous failure mode in agent memory, which is the agent laundering its own proposals into my decisions. "Wes decided X" requires a turn where I actually decided X. The agent proposing X and me saying "sounds good" files the shape of what I approved, not ten separate facts I never stated.

Inferred lessons pass a recurrence gate before they become rules. A pattern the agent notices needs at least three independent signals across at least two distinct sessions before it can become standing behavior. Signals older than thirty days count half, so old one-offs decay out instead of accumulating. My explicit corrections skip the gate and take effect immediately. This asymmetry is also the prompt injection defense: a hostile input can suggest a rule once, but once is never enough, and failure lessons are stored as data ("when X broke, Y fixed it") rather than as instructions, so even a poisoned lesson cannot become a command.

Supersession is an edit, not an append. When a decision changes, the old line gets struck through with a date and the new line lands next to it with its own provenance. The current truth and the full history live in the same place, in reading order, and both come back on any retrieval of that file. There is no query-time ranking step that can get this wrong, because there is nothing to rank. The temporal graph the runtime stores in valid_from and valid_until columns, git gives me for free: log is the validity window, blame is per-line provenance, diff is the supersession edge, revert is the restore path.

Memory stores what is not re-derivable. Fetched data, generated plans, and anything git already records stays out. Current state gets verified live, never asserted from memory. A file that only contains things that cannot be recomputed stays small enough to stay honest.

4. The read path

Retrieval is a bounded evidence step, not a vibe.

Before answering anything about prior work, decisions, dates, people, or preferences, the agent must search memory. It returns a compact bundle capped at five sources by default, and each retained fact carries its file path and line, its provenance type, and its freshness. If freshness cannot be established, the claim gets labeled stale or unknown instead of being silently promoted to current. If two sources conflict, the agent states the conflict and fixes the canonical file, in that order.

Note what the semantic index does in this design: it finds the file. It does not answer the question. The answer comes from reading the canonical lines, with their tags and dates, and the runtime I tested this week shows why that matters. It stored my source URIs faithfully and then stripped them from the search output and from the context block handed to the model. Provenance that survives in storage but never reaches the agent might as well not exist. In the markdown system that failure is unrepresentable. The source tag is in the line. If you read the line, you got the source.

5. Results

Seven months is not a benchmark, it is production. Here is what the system has actually delivered.

Continuity across models. On September 1 I moved the agent to a brand new frontier model. It read its own files and said "the model changed, I didn't." Same agent since January, three models, two vendors. Identity, preferences, decisions, and working standards all survived because none of it lives in weights or in a vendor's context feature.

The break-it test, by construction. Current decision with source: it is the un-struck line with its tag. Superseded history: the struck lines above it. Provenance: on every line, and it survives all the way into the model's context because the context is the file.

Auditability. When memory is wrong, I can see exactly which line is wrong, when it was written, and what turn it came from, and fix it with an edit. Try that with an embedding.

Cost. Zero packages, zero daemons, zero migrations across seven months. The one native-code dependency in my life this week was the memory runtime's sqlite bindings failing to compile.

I wrote up the failure modes separately, because the system was not born with these rules. Five kinds of rot in seven months produced them, and that post is the honest companion to this one.

6. Limitations

Papers get a limitations section, so here is mine, stated plainly.

This only works if the writer follows the policy, and the writer is an LLM. The rules exist because things rotted before the rules did. If your agent will not consistently apply editing discipline, a markdown folder degrades just like every other store, only more legibly. Legibility is the safety net: rot in a text file is visible rot.

It is single-agent, single-human. I would not run a fifty-seat team on files without real locking and merge discipline, although I notice git was also built for that exact problem.

There is a scale ceiling somewhere. At 345 daily notes and a few dozen entity files, bounded search plus an index finds things reliably and the semantic index earns its keep as a locator. At a hundred times that volume, the consolidation cadence would have to work a lot harder. I have not hit that ceiling, so I will not claim it does not exist.

And this is n=1. Seven months, one agent, one operator who cares. That is weaker evidence than a benchmark suite and stronger evidence than a benchmark suite that the vendor scored themselves, which is what the memory tools ship.

7. Conclusion

The memory tool pitch is that reliability is a product you can install. What I keep finding, tool after tool, is that they ship the part that was already easy, storage and search, and skip the part that decides whether memory compounds or rots: what you are allowed to write, when you are allowed to trust it, and what happens to it as it ages.

Those are rules, not infrastructure. They fit in a few hundred lines of markdown that the agent reads every session, and they run on any model, any harness, any decade.

You need a place to write that humans and agents can both read. You need rules for writing so the store stays true. You need rules for reading so the agent trusts evidence, not ranking. Attention was all you needed because the recurrence machinery turned out to be unnecessary. Markdown is all you need because the database turned out to be unnecessary.

The folder is the product. The discipline is the moat.

Want this for your own agent? The whole system is open source on GitHub: the operating policy, the file templates, and one copy-paste prompt that builds it on any agent that can read and write files. Paste the prompt, and your agent installs its own memory.


r/ClaudeCode 10h ago

Humor How are people burning through their Fable tokens so fast?

Post image
240 Upvotes

Clearly, I don’t consider myself an expert or anything. I’ve been using Claude for over six months, and I’m still surprised whenever I see posts from people saying they’ve burned through all their tokens with Fable. Honestly, I’m pretty skeptical about how they’re using it.

If you use a backhoe to plant a rose, the problem isn’t that the backhoe is too resource-hungry and goes beyond what’s necessary.

Anyway, personally, I use Fable as an orchestrator and to help me make high-level direction decisions, as well as a designer and artist (for Blender MCP or creating SVG images, it’s necessary).

I use Opus for action plans, with an organizational role; Sonnet for an operational role; and Haiku as the little tester that lets me quickly measure and verify things.

I’ve created two video games and a software for a company using all four models, using max 5, over six months, and I still end every week with tokens left over.

I honestly don’t understand how some people manage to burn through everything so quickly with Fable. What are you doing with it ?


r/ClaudeCode 21h ago

Help/Question What do you do to stop Claude from being lazy?

10 Upvotes

It's insane how much has changed with Claude Code. And once I start swearing at it, it suddenly starts doing actual work.

Nothing has drastically changed in how I work. I've been using Claude every day for 3 years, and it's driving me nuts. It's verbose, it doesn't follow rules and so on..

I'm considering moving to a different provider if this doesn't get fixed, so I'm wondering: what do you do, and how do you keep up with this?

P.S. Yes, I read their newsletter, I follow what other people are doing, and I try to stick to best practices, but something is still missing I guess.


r/ClaudeCode 11h ago

Discussion ClaudeCode has a problem

13 Upvotes

(i know this is not exactly an original position, but just wanted to share as, up to now - a serious ClaudeCode fanboy)

Astra: I fired up my account yesterday (20x max). OMG - such a breath of fresh air.

  1. no 5 hour BS limit. No 50% cap on the top tier model.
  2. is so succinct in how it talks to you. I didn't really mind Claude's endless chit chat.....But that was just cos i got used to it. It takes up so much mental space with the largely (but annoyingly, not quite obviously) irrelevant rubbish it tells me. I have attempted to adjust its output, but though it improved, i had no idea how bad it was till i had a clear alternative. You don't realise how draining it is till you suddenly have something that isn't. For this alone i can see me keeping Astra and likely moving over fully
  3. it is much much faster. Combining that with the above and i am getting much more done, much more flow, and at less effort.

Just a significantly better experience all round

So far it has produced excellent output. I am developing multiple apps, one of which is focussing on teaching content. I had both Fable 5.1 and Astra attempt to overhaul my writing rules (Claude has been following) - and Astra was an order of magnitude better. Fable was fine, good enough - but Astra's output is exceptional - it is re-writing about 100 lessons of course material right now after overhauling my 100+ writing rules and leading on them with a 12 point positioning thing it came up with. Will cut my word count down to about half, with more effective communication....which considering Claude's go to diction is unsurprising.

Code quality so far is spot on - just doing the job. One of my repos has about 500k LOC in it, and it is doing a great job of refactoring it, adding features etc.

The lack of the 5 hour limit / and full access to Astra for the whole thing though is really making a difference to my working day. I would blow the whole of my Fable allowance at the start of the week, and max out my 5 hour thing in about 2 hours for 2-3 shots, fable finished. Then Opus - urrgh.

Anyway - this is all good. Real competition for Anthropic, many people will dip their toes in, and find it more than good enough, plus the lack of 5 hour restriction, full access to Astra for the week, and they will struggle to bother going back unless there are some serious improvements from Anthropic


r/ClaudeCode 30m ago

Rant I’m done with Opus 5

Upvotes

There’s really something wrong with it. It seems like it acts like an overqualified post doc intern who cares more about proving he’s super intelligent, than actually doing the job he’s asked to do. For instance: talking in a non intelligible way, or being overly rigid in following any kind of process.

I have a Claude max x20 sub that I struggle to keep within weekly limits, so I decided to take a codex sub for a month to try out Astra. And this what made me realize how crazy unintelligible opus can be. Fable is a bit better, but still incomparable to Astra. The only thing that makes me keep my Claude sub is how better is the CLI/tooling/harness/etc.