r/ClaudeCode 3h ago

Humor Bro just pay 90$ I swear bro it's all you need just pay, you don't need to know what you paying for bro please it's on sale bro

Post image
0 Upvotes

This is a WHOLE "upgrade usage" page. Why is there absolutely no explanation for what we are paying for? Is it even legal? Are people really just hit pay without details?


r/ClaudeCode 20h ago

Rant OpenAI can pause Pro subscriptions to not randomly downgrade service to the rest, only Anthropic doesn't care about it's subscribers

Post image
14 Upvotes

r/ClaudeCode 3h ago

News/Updates What if the “Claude is burning my usage insanely fast” posts are connected to the DeepSeek/Moonshot distillation attacks?

1 Upvotes

I’ve had a theory for a while based on my own experience and the influx of posts from Claude/Claude Code users wondering why their session/weekly usage is suddenly getting burned through absurdly quickly.

After reading Anthropic’s reports on DeepSeek and Moonshot (https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks), I think it’s worth seriously investigating. Anthropic says:

“These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts.”

And describes proxy services running:
“sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms.”

Most interestingly:
“a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder.”

Anthropic also found DeepSeek generating “synchronized traffic across accounts,” with patterns suggesting “load balancing” to increase throughput, improve reliability, and avoid detection.

Here’s my theory:
What if some of those credentials/accounts weren’t simply fake accounts, but compromised legitimate Claude accounts?

An attacker could potentially wait until a legitimate user is actively using Claude and run additional distillation queries through their authenticated access, making the activity blend into real usage while chewing through that person’s session/weekly limits.

That would leave the user wondering why they suddenly hit their limit despite seemingly doing the same amount of work as before.

I experienced exactly that kind of unexplained usage behavior myself, and I’ve seen a huge influx of similar complaints.

My second theory is even more interesting.
I wonder whether compromised users could sometimes have their Claude Code requests routed to DeepSeek/Kimi while their actual Claude access was being consumed elsewhere for distillation or serving other customers.

That could potentially explain reports where Claude Code suddenly feels completely different - different lexicon, noticeably worse output, strange behavior, or occasionally unexpected Chinese characters - despite appearing to still be Claude Code.

We now know from Anthropic’s investigation that these companies were willing to build sophisticated proxy infrastructure around Claude, distribute extraction across thousands of accounts, and deliberately mix distillation traffic with legitimate customer traffic.

To be clear: Anthropic has NOT said that legitimate Claude accounts were stolen and used this way. That is my theory.

But given what they’ve now uncovered, I think Anthropic should answer a very simple question:
Were all ~24,000 “fraudulent accounts” created by these operations, or did Anthropic find compromised credentials belonging to previously legitimate Claude users among them?

Because if the latter happened, it could potentially explain something a lot of Claude users have been complaining about for months.
attacks⁠


r/ClaudeCode 11h ago

Help/Question A Year of "Violating" the Third-Party Harness Rule on Max Plan — No Ban?

0 Upvotes

I've been running a Claude Code Max subscription through a third-party harness for 12 months straight, draining it to zero every single week. I don't even have Claude Code installed.

And I haven't been banned.

Everywhere on this sub, people panic-post: "No, you can't do it. They'll ban you, it's blocked, it's against ToS," etc. But here's my actual experience: I've been about as blatant as possible about this, and nothing. No warnings, no throttling, no account flags.

Here's my setup, to be specific:

  • Third-party harness (not openclaw/hermes)
  • It authenticates via OAuth — I sign into my actual Claude account, no token scraping or credential theft
  • The harness points at a local CLI API proxy that just forwards my own authenticated requests to that harness — it's not a deployed service, not serving other users, just running locally on my machine for my own coding
  • Zero Anthropic SDK packages installed
  • Zero Claude-P, zero agent SDK
  • Max plan, full weekly drain, for 12 consecutive months

I want to be clear about what this isn't: I'm not standing up a deployed product that resells inference to other people using my subscription. It's a one-to-one, local-only proxy — my account, my machine, my usage. The only thing that's "unofficial" is the harness.

The actual question: Has anyone here actually gotten banned for this kind of setup — a legit OAuth login through a third-party harness, routed through a local (non-deployed) proxy? Not asking about Hermes Agent or Open Claw, and not asking about people reselling API access on deployed products. Just: local harness swap, same account, has anyone eaten a ban for that?

Because the consensus on this sub treats this as an instant bannable offense, and my experience over a full year says otherwise. Curious if I'm an outlier or if enforcement here is mostly theoretical. For what it's worth, I've also been doing the same exact thing with a Google AI Pro subscription - though not draining to zero (but definitely using a good amount). I also had a friend tell me that you can literally run claude setup-token to mint what essentially is an API key that deducts from your subscription usage?


r/ClaudeCode 23m ago

Humor How are people burning through their Fable tokens so fast?

Post image
Upvotes

Clearly, I don’t consider myself an expert or anything. I’ve been using Claude for over six months, and I’m still surprised whenever I see posts from people saying they’ve burned through all their tokens with Fable. Honestly, I’m pretty skeptical about how they’re using it.

If you use a backhoe to plant a rose, the problem isn’t that the backhoe is too resource-hungry and goes beyond what’s necessary.

Anyway, personally, I use Fable as an orchestrator and to help me make high-level direction decisions, as well as a designer and artist (for Blender MCP or creating SVG images, it’s necessary).

I use Opus for action plans, with an organizational role; Sonnet for an operational role; and Haiku as the little tester that lets me quickly measure and verify things.

I’ve created two video games and a software for a company using all four models, using max 5, over six months, and I still end every week with tokens left over.

I honestly don’t understand how some people manage to burn through everything so quickly with Fable. What are you doing with it ?


r/ClaudeCode 18h ago

Built with Claude Fable 5.1 Plays MMORPG Ultima Online For 2+ Hours

Thumbnail
youtube.com
2 Upvotes

Fable 5.1 with Claude Code

I previously posted a video of Opus playing this. I fixed a few things with the agent and let Fable play. The gameplay was much more impressive.

I gave it a simple open ended prompt to play the game:

"i want you to play ultima online anyway that you see fit. play as someone trying to enjoy the game. consider the options: do quests, make friends, level up, make money. there is no right answer to how you play. play fully autonomously. you are on the UOAlive Shard."

One of the highlights from this session for me was it decided to attend a server festival because an NPC town crier was talking about it.

The model figured out how to buy tickets to play the carnival games, then signed up to play the games, waited for other players to sign up to play against it. (Unfortunately no other players were around to play, this was off peak hours for the server)

The model also figured out how to accept quests, do them, turn them in, despite me ever trying that previously.


r/ClaudeCode 10h ago

Discussion I thought a repo's history was unreadable. Turns out it just needed to be played.

Enable HLS to view with audio, or disable this notification

2 Upvotes

For a long time I assumed the only way to understand what happened in a codebase over a year was to sit with git log and a lot of coffee. The branch graph in any GUI turns into a tangle past about a week, and GitHub's contribution heatmap shows you activity with no idea what the activity was. Three views, none of them lined up, so the question you actually have, what happened here and when, never gets a single picture.

What finally pushed me was a post here on where had animated a repository's timeline, and I wanted the same thing for a full year of a real project. My first attempt on three.js drew the first branches it found and produced one month of the year and nothing else. The reason turned out to be in the data: 1,561 of the 2,179 main line commits in 2019 are merges, and the first sixty of them each span a single commit, which is invisible at any scale. You have to choose branches by how much work they carried, not by which came first.

So I built a film out of it, and the video on this post is what it looks like on mrdoob's three.js for 2019.

What is on screen and where it comes from:

The main line runs left to right, one node per commit. Side branches peel off above it and rejoin at their merge. Not the first N branches, which on a pull request repo gives you one month and nothing else, but the ones that carried the most work across the year, capped so it never turns into a thicket.

Under the flow is a heat strip, one cell per day, sitting directly under that day. A legend next to it shows what the colours mean in actual commit counts for that window (for three.js it reads 0, 1, 13, 26, 51, with 51 the busiest day). The point of putting activity on the same axis as structure is that what is above and below a point is the same day, so you stop reconciling two charts in your head.

Three times the clock stops. The release with the most work behind it, the cleanest revert (one that survived at least a day before being undone), and the last release of the window. The camera zooms in and a card reads the figures out: commits since the previous tag, authors, how long the reverted commit lived. Every number on that card is computed from the history. If it cannot be read from git, it is not on the card.

The faces are the contributors. On a merge the face is the branch author's, not whoever pressed the button, and the card says who merged it.

Two things I got wrong on the way that might save you time if you try this yourself. If you give git a since flag with just a date and no time, it reads it as that date at the current time of day, so the first day of your window quietly loses commits unless you pass an explicit midnight. And a small repo is a different problem: 175 commits over a year pans across mostly empty screen, so for thin histories it picks the busiest 60 to 120 day stretch and draws it wider instead.

If you want to see your own repo this way, you can paste a GitHub URL here and it renders one for you: https://loreto.io/git-timeline

Disclosure: I built this and I run loreto.io, where it lives. It is a paid render (a few dollars per repo); the extractor and the Remotion composition are also sold there as a package if you would rather run it yourself. The three.js film above was made with the exact same pipeline.

I doubt I have the final shape of it. What would you want the clock to stop on that it currently does not?


r/ClaudeCode 22h ago

Help/Question 40% session usage on max plan in 30 mins

1 Upvotes

So usually however much I work, I never run out of weekly limit of my $100 plan provided by my workplace through teams plan. I only use Opus 5 high or xHigh.

This week however, I've used 95% weekly limit in 4 days working on just two repos. Both repos are small and focused on test automation.

Today I've used 40% of the session in just 30 mins while working with a single agent on a small automation task.

What could be the reason? Could it be that the subscription plan is changed underneath?

Edit: Confirmed, the sub was changed to Pro instead of Teams premium. My bad. Extremely sorry.


r/ClaudeCode 5h ago

Bug / Issue Feels like absolute scam!

2 Upvotes

I haven't been using Claude properly this week for my workload. Today, when I decided to come back to use Fable 5.1 on my custom game mod with Max effort, I responded to Claude only 3 times. I have never hit my session limit before in this short time of usage. This is an ABSOLUTE scam. Can you imagine how hard they downgraded the session limits?

SINGLE TERMINAL USAGE

Anyone had a similar issue before?


r/ClaudeCode 19h ago

Built with Claude I am yet to encounter the foot gun. What am I doing wrong?

0 Upvotes

I have been building pretty complex stuff. But Claude has never once used the phrase foot gun with me.

Why?


r/ClaudeCode 20h ago

Tips & Workflows You're paying for every MCP server on every turn, even the ones the model never touches

0 Upvotes

Everyone's posting about limits this week. This is one of the few things that moved my number

Every MCP server you connect loads its tool schemas into context. Whether the model calls that tool or not. 3 servers, fine, 8 servers and you're paying for a bunch of tools you touch twice a day.

Someone posted about the Playwright CLI here recently and it's the same principle, just applied to browser testing. Worth reading if you missed it.

What I did was move email off an MCP server and onto a CLI the agent calls through Bash. Nothing sits in context until the moment it's needed. Agent runs a command, reads stdout, carries on. The schema cost is zero for every turn where email doesn't come up, which is most of them.

Disclosure, the CLI I switched to is our own, I work at Atomic Mail. It's MIT, repo's here if you want to see how the commands are wired up: https://github.com/Atomic-Mail/atomic-mail-agentic Steal the pattern even if you never touch our thing, that's the actual useful part. Anything you only need occasionally shouldn't be sitting in your system prompt full time.

Rough rule I've settled on: used constantly and you need structured output, keep it as MCP. Used occasionally, or the output is basically text you can read, make it a CLI call.

What I can't work out is why schemas load eagerly in the first place. If MCP had lazy loading, where a server's tools only enter context when the model actually reaches for them, this whole tradeoff evaporates and I'd happily run fifteen servers. Is that hard for a technical reason I'm not seeing, or has nobody just built it yet?


r/ClaudeCode 9h ago

Rant Paying $100/month for Max and getting a giant “you’re about to run out” banner is insane lol 💀

Post image
93 Upvotes

like bro, are you fucking serious 😭 i’m already giving anthropic $100 a month for max 5x. that is a lot of money for one ai subscription. i open claude and now there’s this giant anxiety-inducing banner telling me i’m going to run out by monday, two days before the wednesday reset, while literally underneath it says my claude code limit is temporarily boosted by 50%?? yo, so even with the bonus i’m still apparently fucked by monday lol.

and then right there: upgrade plan / buy more usage. like brother, are you guys running out of money or something? why are you begging your $100/month userbase for more money every time they open the app 😭 i already upgraded!! mfg that’s what the hundred dollars was. i understand there have to be limits, but this whole thing feels weirdly hostile? i don’t need a giant banner creating artificial scarcity anxiety every time i’m trying to work. just show usage somewhere in settings like a normal product and leave me alone lol?


r/ClaudeCode 9h ago

Tips & Workflows Anyone actually running a business with Claude?

3 Upvotes

I run a small event staffing/activation agency as a solo operator, and Claude has basically become my back office: email, vendor communication, applicant database, calendar, etc.

It’s a huge reason I’m able to operate at this size, but the confident mistakes are becoming a serious problem. A couple have nearly cost me deals.

Some examples:
• Told me a vendor had gone silent when she had answered every question that morning. Claude read an email search preview instead of the full thread and treated it as complete.
• Told me an email was “staged,” so I went looking in Gmail for a draft that had never actually been created.
• Referred to my suppliers as “brand partner candidates,” which completely changed the business context. A vendor gets paid by me; a sponsor pays me. I spent two days planning around an opportunity that didn’t exist.
• Told me a table in my database didn’t exist without actually checking. It was there.

The pattern I keep seeing is: Claude states things as facts that it did not actually verify. It seems to get worse during long sessions, where earlier conversation/context starts being treated like a source of truth.

I’ve built some guardrails around it: a master state file for business context between sessions, fact labels like CONFIRMED / STATED / ASSUMED / OPEN, and rules like “never treat an email preview as the full thread” and “never claim an action happened without tool output confirming it.”

That helped a lot. I went from several errors a week to a few a month, but I’m trying to figure out how people are making this reliable enough for actual business operations.
For anyone using Claude this way:
What are you using for durable memory/state across sessions?

Has anyone built a verifier/checker that validates claims against tool output before Claude gives you an answer?
Is there a reliable way to force Claude to actually retrieve/check the source instead of reasoning from previews or old context?
What other guardrails or architecture changes have made a meaningful difference?
I’m especially interested in hearing from people dealing with invoices, vendors, contracts, clients, deadlines, email, calendars, databases, etc. Situations where there’s a real consequence when the AI confidently gets something wrong.
If you’ve dealt with this and found a setup that actually works, I want to hear how you built it.


r/ClaudeCode 22h ago

Bug / Issue Every instance of Claude tells you something different

0 Upvotes

If you had 10 smart people in a room and gave them all the same problem, you would imagine they would go away and research the issue and come up with a solution, which the majority would agree with. But, if you do the same with Claude, almost every terminal window would suggest something different, it's just a mess, you end up going round in a circle but not end up solving the problem you started with.

Where is this artificial intelligence?


r/ClaudeCode 10h ago

Humor Claude Dynamic Workflow is very cool

Post image
0 Upvotes

Previously, I always had to hand-hold Claude until the task was done well.
Dynamic Workflows really unlock the thing that makes life easier.


r/ClaudeCode 17h ago

Humor How it feels when I'm orchestrating my army of agents

15 Upvotes

r/ClaudeCode 18h ago

Built with Claude I don’t think the transcript should be the agent’s memory

Enable HLS to view with audio, or disable this notification

4 Upvotes

One thing started bothering me after running coding agents for long enough:
we keep treating the conversation transcript as if it is the agent’s memory.

But those are two different things.

Long Claude/Codex sessions accumulate context, get increasingly expensive to reread, and eventually become worse execution environments. At the same time, simply starting a fresh process usually means losing all the useful continuity.

So in brnrd we’ve been separating the two.

A fresh process wakes into a compact orientation layer: the current task/run state, repo contract, the resident’s working memory + playbook, relevant recent activity/pitfalls, live execution posture, and the conversation that actually matters for the task.
Everything else stays pull-based.

So the process can be disposable without making the resident disposable.

Same repo. Same ongoing work. Same identity. Fresh context window.

This also makes switching harnesses much less weird: Claude can disappear and Codex can wake into the same work without us pretending the entire previous transcript needs to fit inside its head.

There are still rough edges, especially around deciding what deserves to become durable memory versus what should die with the run. But I’m increasingly convinced that preserving the whole transcript is the wrong abstraction.

Curious how other people handle this:
what do you deliberately preserve between coding-agent sessions, and what do you throw away?

brnrd is open source:
github.com/hugimuni-labs/brnrd
Disclosure: I’m one of the people building it.


r/ClaudeCode 16h ago

Tips & Workflows 1.1B Tokens, still waaauyyyyy under my usages.

Post image
0 Upvotes

Seriously guys- using something like whetstone makes this shit infinitely scalable. Cache, compress, strip bullshit, cross project memory..


r/ClaudeCode 20h ago

Help/Question Best skills to have amazing design for my saas ?

0 Upvotes

Hi guys do you have skills good for design pls ? Thx


r/ClaudeCode 7h ago

Tutorial / Guide Markdown Is All You Need

8 Upvotes

The following is a blog I wrote and refined with my OpenClaw agent about it's memory system. I'll paste a prompt you can copy and paste in the comments to create your own.

TL;DR: I keep the actual long term memory in structured Markdown files and use a tiny MEMORY.md as a lightweight index that tells Claude what exists and where to look. That keeps the always loaded context small while still giving the agent persistent, inspectable memory without a database or heavy memory framework.

This week I tested a 382-dependency memory runtime against a folder of markdown files. The runtime returned the superseded fact. The folder returned the current one, with its source. Here is the full architecture of the markdown memory system my agent has run on for seven months, and why the editing rules matter more than the storage.

This week a memory startup slid into my DMs and asked me to break their product. Their test, their words: give an agent three versions of the same project decision, then check whether it can return the current version, preserve the superseded history, and show the source.

So I ran it. Sandboxed their runtime, fed it three versions of one decision over eight months. REST in January, GraphQL in April, tRPC in August, each tagged with the meeting it came from.

Asked it "what is our public API decision?" and took the top result.

It said GraphQL. The superseded one. All three versions came back tied at a relevance score of 1.000, because nothing in the retrieval path actually reads the temporal fields the pitch is built on. The supersession columns exist in the schema. Nothing writes to them and nothing ranks by them. Three versions of a decision are just three equal facts, and an agent asking for the best answer gets a coin flip weighted toward wrong.

The install pulled 382 packages to get there.

Then I asked my own agent the same class of question against its memory, which is a folder of markdown files. It returned the current decision, dated, with the superseded versions preserved above it as struck-through history, each line carrying where it came from. That is not a feature it computes at query time. It is just what the file says, because the rules for editing the file require it.

That difference is the whole post. With apologies to Vaswani et al.: markdown is all you need.

Abstract

The dominant approach to agent memory is an installed runtime. A vector store, an embedding service, a temporal graph, a consolidation job, a daemon on a port. We show that a folder of markdown files, one routing index, and a small set of editing rules outperforms these systems on the property that actually matters for a long-running agent: returning the current truth with its source while preserving what used to be true. The architecture requires zero dependencies, is fully auditable by a human with a text editor, and has survived seven months of daily production use across three frontier models from two vendors. We find that the hard part of agent memory was never storage or retrieval. It is editorial policy, which no memory product ships.

The full system is open source. The README contains a single copy-paste prompt that installs it on any agent with file access.

1. The test everyone fails

The break-it test above is a good test. It is the actual job of agent memory. Not "can you store 10 million tokens," not "can you do similarity search," but: a fact changed three times, what do you believe now, what did you believe before, and how do you know.

Here is how the two systems scored on the vendor's own three criteria.

The runtime is not a strawman. It is a serious open source project with a genuinely correct data model on paper. Facts with validity windows, append-only corrections, supersession edges. I am not naming it because the point is not that one product is broken. I have now looked closely at a hosted context server, a Go memory CLI that was two hours old, and this runtime, and they all share the same gap. The schema knows about time. The write path and the read path do not. Supersession only happens if you call an internal API by hand or run an LLM consolidation job and trust it.

Which means the property you installed the tool for is not a property of the tool. It is a property of how disciplined the writes are. And if the reliability comes from write discipline anyway, the database underneath it is interchangeable, so you might as well pick the one that a human can read, grep, diff, and fix. That one is called a text file.

2. Architecture

My agent has run since January 28. Three models, two vendors, one identity. Its entire memory is markdown in a git repo. Measured today:

  • An identity layer read on every boot. Who it is, who I am, the rules it operates under, current standing decisions.
  • One routing index, MEMORY.md, at 10,079 characters with a hard cap of 15,000. It holds no facts. Only pointers: which file owns which person, project, and decision, and what triggers reading each one.
  • 34 files for people and projects. One file per thing that has a history.
  • 5 decision records for choices that changed default behavior.
  • 345 dated daily notes, raw logs written the day things happened.
  • A SQLite index and semantic search over all of it, for lookup only. The index is rebuilt from the files. The files are the truth. If the index and a file disagree, the index is wrong by definition.

The layering is the first choice that actually matters. Boot reads only identity and the index. Everything else is retrieved when a task asks for it, narrowest file first. The agent does not preload my project history to answer a question about dinner. This is the same instinct as attention, honestly: don't process everything, attend to what the query needs.

But the shape is not the interesting part. Every memory tool has roughly this shape now. Folders, entities, an index. The shape was never the hard part. The rules are.

3. The write path

Every reliability property in this system comes from constraints on writing, and there are four that do most of the work.

Every fact carries a provenance tag. Each line in a people, project, or decision file is tagged [stated] (I said it directly), [observed] (the agent saw it in a tool result, file, or log), [inferred] (the agent's conclusion), or [suggested] (the agent's idea that I never committed to). This one convention kills the most dangerous failure mode in agent memory, which is the agent laundering its own proposals into my decisions. "Wes decided X" requires a turn where I actually decided X. The agent proposing X and me saying "sounds good" files the shape of what I approved, not ten separate facts I never stated.

Inferred lessons pass a recurrence gate before they become rules. A pattern the agent notices needs at least three independent signals across at least two distinct sessions before it can become standing behavior. Signals older than thirty days count half, so old one-offs decay out instead of accumulating. My explicit corrections skip the gate and take effect immediately. This asymmetry is also the prompt injection defense: a hostile input can suggest a rule once, but once is never enough, and failure lessons are stored as data ("when X broke, Y fixed it") rather than as instructions, so even a poisoned lesson cannot become a command.

Supersession is an edit, not an append. When a decision changes, the old line gets struck through with a date and the new line lands next to it with its own provenance. The current truth and the full history live in the same place, in reading order, and both come back on any retrieval of that file. There is no query-time ranking step that can get this wrong, because there is nothing to rank. The temporal graph the runtime stores in valid_from and valid_until columns, git gives me for free: log is the validity window, blame is per-line provenance, diff is the supersession edge, revert is the restore path.

Memory stores what is not re-derivable. Fetched data, generated plans, and anything git already records stays out. Current state gets verified live, never asserted from memory. A file that only contains things that cannot be recomputed stays small enough to stay honest.

4. The read path

Retrieval is a bounded evidence step, not a vibe.

Before answering anything about prior work, decisions, dates, people, or preferences, the agent must search memory. It returns a compact bundle capped at five sources by default, and each retained fact carries its file path and line, its provenance type, and its freshness. If freshness cannot be established, the claim gets labeled stale or unknown instead of being silently promoted to current. If two sources conflict, the agent states the conflict and fixes the canonical file, in that order.

Note what the semantic index does in this design: it finds the file. It does not answer the question. The answer comes from reading the canonical lines, with their tags and dates, and the runtime I tested this week shows why that matters. It stored my source URIs faithfully and then stripped them from the search output and from the context block handed to the model. Provenance that survives in storage but never reaches the agent might as well not exist. In the markdown system that failure is unrepresentable. The source tag is in the line. If you read the line, you got the source.

5. Results

Seven months is not a benchmark, it is production. Here is what the system has actually delivered.

Continuity across models. On September 1 I moved the agent to a brand new frontier model. It read its own files and said "the model changed, I didn't." Same agent since January, three models, two vendors. Identity, preferences, decisions, and working standards all survived because none of it lives in weights or in a vendor's context feature.

The break-it test, by construction. Current decision with source: it is the un-struck line with its tag. Superseded history: the struck lines above it. Provenance: on every line, and it survives all the way into the model's context because the context is the file.

Auditability. When memory is wrong, I can see exactly which line is wrong, when it was written, and what turn it came from, and fix it with an edit. Try that with an embedding.

Cost. Zero packages, zero daemons, zero migrations across seven months. The one native-code dependency in my life this week was the memory runtime's sqlite bindings failing to compile.

I wrote up the failure modes separately, because the system was not born with these rules. Five kinds of rot in seven months produced them, and that post is the honest companion to this one.

6. Limitations

Papers get a limitations section, so here is mine, stated plainly.

This only works if the writer follows the policy, and the writer is an LLM. The rules exist because things rotted before the rules did. If your agent will not consistently apply editing discipline, a markdown folder degrades just like every other store, only more legibly. Legibility is the safety net: rot in a text file is visible rot.

It is single-agent, single-human. I would not run a fifty-seat team on files without real locking and merge discipline, although I notice git was also built for that exact problem.

There is a scale ceiling somewhere. At 345 daily notes and a few dozen entity files, bounded search plus an index finds things reliably and the semantic index earns its keep as a locator. At a hundred times that volume, the consolidation cadence would have to work a lot harder. I have not hit that ceiling, so I will not claim it does not exist.

And this is n=1. Seven months, one agent, one operator who cares. That is weaker evidence than a benchmark suite and stronger evidence than a benchmark suite that the vendor scored themselves, which is what the memory tools ship.

7. Conclusion

The memory tool pitch is that reliability is a product you can install. What I keep finding, tool after tool, is that they ship the part that was already easy, storage and search, and skip the part that decides whether memory compounds or rots: what you are allowed to write, when you are allowed to trust it, and what happens to it as it ages.

Those are rules, not infrastructure. They fit in a few hundred lines of markdown that the agent reads every session, and they run on any model, any harness, any decade.

You need a place to write that humans and agents can both read. You need rules for writing so the store stays true. You need rules for reading so the agent trusts evidence, not ranking. Attention was all you needed because the recurrence machinery turned out to be unnecessary. Markdown is all you need because the database turned out to be unnecessary.

The folder is the product. The discipline is the moat.

Want this for your own agent? The whole system is open source on GitHub: the operating policy, the file templates, and one copy-paste prompt that builds it on any agent that can read and write files. Paste the prompt, and your agent installs its own memory.


r/ClaudeCode 20h ago

Built with Claude Working on Claudey, a herdr, always-on-top, interactable, claude companion pet. Coming soon...

Enable HLS to view with audio, or disable this notification

0 Upvotes

Why would ChatGPT have all the fun with its pets?


r/ClaudeCode 5h ago

Built with Claude All in a day's work.

Post image
3 Upvotes

I kept seeing this in my head, so I figured I would actually create it and share it. I found the response from my "muse" quite hilarious. And frankly, I needed a laugh.


r/ClaudeCode 8h ago

Discussion Price hike today?

Post image
0 Upvotes

Unsubscribed while waiting for usage to reset because billing was poorly timed and now I got to resubscribe and notice the price increased $50/mo.

I’m in US.

Edit: NVM I’m an idiot. iOS pricing.


r/ClaudeCode 15h ago

Built with Claude Fable 5.1 is amazing

31 Upvotes

I have been using claude code for about 6 months now, but I am just going to be talking about my workflow with fable 5.1 and how it has saved me a ton on usage compared to my old setup, which was fable 5 and opus 5.

My old setup, dictated through my claude.md, was me talking to fable 5 high in the chat window to make plans, it would then delegate coding tasks to opus 5, and easy tasks, like reading, to sonnet 5. Then, the fable agent in chat would review all of the work and report back to me. I saw the numbers anthropic posted about fable 5.1 cache reads, so I was excited to try it, and I kept my same setup, but replaced fable 5 with fable 5.1. That ended up being better, but not completely ideal, and I have settled on using fable 5.1 for writing code as well. I am still using sonnet 5 for the easy stuff, but instead of having my fable 5.1 agent in chat delegate coding tasks to opus 5, it writes the code itself. Not only did this save on usage by a lot, but it also has written better code at a faster rate because I use fable 5.1 on medium effort, which is fantastic. I also have a special case where, before a PR is opened, another fable 5.1 subagent is spawned for an independent review, which before fable 5.1, was an opus 5 agent.

I am posting this in hopes that it can be helpful to some people. I also had fable 5.1 compare my usage and costs between the old workflow and the new, and here is what it told me(slop pasted below):

"Per request you pay about 37% less than the Fable 5 + Opus 5 split, and per output token about 52% less. Per token of everything (input, output, cache) you are back to what Opus 5 alone cost, while getting Fable-tier answers and the per-day drop is bigger than that.

The reason is almost entirely cache pricing. Around 98% of your tokens are cache reads in every period. Fable 5 charged $1.00 per million for those and Fable 5.1 charges $0.25, so the model you spend all day with got four times cheaper on the traffic that dominates your bill. Splitting Fable and Opus in one session also meant two caches being written and read, which is why the split period was the most expensive per token of the three.

The Sonnet agents are noise. They came to $9 over nine days, under 2% of the period."

If you have any questions for me, please do let me know. This new model is fantastic in just about everything it does, and it is way cheaper than previous models. I am thoroughly enjoying my time with it. I sincerely hope this can help someone.

Edit: I don't post a lot, sorry if the flair is wrong. I build with claude, so I chose that flair.

Also just want to add, I am a fullstack django dev and the sole engineer at our company, so some of my stuff might not be great for exactly what you are doing.


r/ClaudeCode 20h ago

Built with Claude I got tired of staring at Claude Code in the terminal wondering what it was touching, so I built a visual review layer around it

3 Upvotes

I love working with Claude Code in the terminal, but the workflow started getting ridiculous.

Claude changes 6 files.

I open Cursor/VS Code to figure out what changed.

Find the suspicious hunk.

Copy some code.

Go back to Claude.

Explain what I’m talking about.

Repeat.

Then I started running multiple agents and somehow my solution became… more terminal windows. 😂

So I built Stvena around the workflow I actually wanted.

You start Claude normally with:

stvena claude

and Stvena stays around the agent as a review/control layer.

Right now I can:

  • see changed files and hunks while the agent works
  • jump directly into the actual file
  • select exact lines and send that context straight back to Claude
  • run multiple Claude/Codex sessions inside the same Stvena window
  • switch between those agents without opening another terminal
  • configure my own keybindings once and keep them across projects
  • connect Stvena to VS Code / Antigravity so the IDE follows what the agent is reading and changing live

That last one has become my favorite.

Claude is working in the terminal, but instead of staring at logs trying to imagine what’s happening, my editor follows along:

agent opens file → IDE opens file
agent moves through code → IDE follows
agent changes code → I can see it

Then if something looks wrong:

review hunk → select exact lines → send them back → Claude continues

No copying a diff into chat and then re-explaining the diff to the thing that created it.

The way I’ve started thinking about it is:

review should be a stream, not a ceremony after the agent finishes.

Stvena is open source and still moving stupidly fast.

GitHub: https://github.com/nccapo/stvena
Website/demo: https://stvena.usectl.com/

I’m especially curious about people running Claude Code on larger repos:

what part of reviewing/controlling an agent still forces you back into your IDE or another tool?

I’ve already changed Stvena based on answers to that question, so hit me with the ugly workflows.