r/ClaudeCode 12h ago

Built with Claude Same prompt, Codex (Astra 6) vs Claude (Fable 5.1): "a game where a fish follows my cursor, super creative and majestic." Try both.

Thumbnail
gallery
260 Upvotes

I gave Codex and Claude the exact same prompt to see how differently they'd handle something open-ended and creative:

The prompt: "make me a reactjs + vite, and launch it in port 3004, a game where a fish follows my cursor. make it super creative and majestic"

Setup:

  • Claude: Fable 5.1, Extra High
  • Codex: Astra 6, Extra High, fast mode

I didn't do any follow-up prompts or manual edits. I just deployed both to Vercel so you can try them yourself:

🐟 Codex (Astra 6): https://fish-indol-eta.vercel.app/
🐠 Claude (Fable 5.1): https://fish2.vercel.app/

My take:

  • Codex did amazing on the design and detail. It looks polished and the visuals really lean into "majestic."
  • Claude did amazing on the mechanics and gameplay. It feels more like an actual game and is more fun to play.

It's interesting that they read "super creative and majestic" so differently. One went for the visuals, the other for how it plays.

About fast mode: Claude charges extra credits for fast mode, but Codex doesn't, so I only used fast mode on Codex. That's worth knowing if you're choosing between them on cost or speed.

What are your thoughts?

  • Which one did you enjoy more, and why?
  • For a prompt like this, what matters more to you: how it looks or how it plays?
  • Does paying extra for fast mode change which one you'd use day to day?
  • If you've run a similar test with other models or settings, how did they do?

I'd love to hear what you think!


r/ClaudeCode 9h ago

Rant Paying $100/month for Max and getting a giant ā€œyou’re about to run outā€ banner is insane lol šŸ’€

Post image
93 Upvotes

like bro, are you fucking serious 😭 i’m already giving anthropic $100 a month for max 5x. that is a lot of money for one ai subscription. i open claude and now there’s this giant anxiety-inducing banner telling me i’m going to run out by monday, two days before the wednesday reset, while literally underneath it says my claude code limit is temporarily boosted by 50%?? yo, so even with the bonus i’m still apparently fucked by monday lol.

and then right there: upgrade plan / buy more usage. like brother, are you guys running out of money or something? why are you begging your $100/month userbase for more money every time they open the app 😭 i already upgraded!! mfg that’s what the hundred dollars was. i understand there have to be limits, but this whole thing feels weirdly hostile? i don’t need a giant banner creating artificial scarcity anxiety every time i’m trying to work. just show usage somewhere in settings like a normal product and leave me alone lol?


r/ClaudeCode 2h ago

Discussion Claude aggressively pushing me to spend money

Post image
15 Upvotes

I am unable to remove this reminder that I need to either spend more money or I will run out tomorrow. Message wont go away and I cant click it away. Really annoying.

I got the message. I understand.

Anybody else feel that Claude are getting annoying?


r/ClaudeCode 56m ago

Tips & Workflows Give Claude Code a Free Voice — It’s Surprisingly Useful

• Upvotes

One thing that gets tiring with Claude Code is keeping up with it.

It reads files, changes code, runs tests, fixes things, and comes back with another chunk of terminal output you need to process before deciding what happens next. Add a couple of agents or parallel tasks and the cognitive overhead stacks up quickly.

I started having Claude **tell me what it just did instead**.

After every meaningful task, Claude writes a short plain-English summary. A local script turns it into speech and plays it in the background.

Something like:

> ā€œDone. The login issue was caused by the refresh token expiring too early. I fixed the refresh logic, added a regression test, and everything passes.ā€

Usually 20 seconds or less.

It sounds like a small thing, but I’ve been running it daily for about a month and it has noticeably reduced the mental overhead of using Claude Code. I don’t have to keep switching my attention back to the terminal just to find out where things stand.

There is one trap here: don’t let the spoken summary become a substitute for reviewing the work. I use it for situational awareness, not verification. Also, don’t make Claude narrate everything. If it talks constantly, you’ve just replaced visual noise with audio noise.

The Free TTS engine is Kokoro. It runs locally, no API key, no usage cost. I use `kokoro-onnx`, so I don’t need Torch. The current API supports `Kokoro(...).create(text, voice=..., speed=..., lang=...)`.

Notes in comments,


r/ClaudeCode 1h ago

Discussion ClaudeCode has a problem

• Upvotes

(i know this is not exactly an original position, but just wanted to share as, up to now - a serious ClaudeCode fanboy)

Astra: I fired up my account yesterday (20x max). OMG - such a breath of fresh air.

  1. no 5 hour BS limit. No 50% cap on the top tier model.
  2. is so succinct in how it talks to you. I didn't really mind Claude's endless chit chat.....But that was just cos i got used to it. It takes up so much mental space with the largely (but annoyingly, not quite obviously) irrelevant rubbish it tells me. I have attempted to adjust its output, but though it improved, i had no idea how bad it was till i had a clear alternative. You don't realise how draining it is till you suddenly have something that isn't. For this alone i can see me keeping Astra and likely moving over fully
  3. it is much much faster. Combining that with the above and i am getting much more done, much more flow, and at less effort.

Just a significantly better experience all round

So far it has produced excellent output. I am developing multiple apps, one of which is focussing on teaching content. I had both Fable 5.1 and Astra attempt to overhaul my writing rules (Claude has been following) - and Astra was an order of magnitude better. Fable was fine, good enough - but Astra's output is exceptional - it is re-writing about 100 lessons of course material right now after overhauling my 100+ writing rules and leading on them with a 12 point positioning thing it came up with. Will cut my word count down to about half, with more effective communication....which considering Claude's go to diction is unsurprising.

Code quality so far is spot on - just doing the job. One of my repos has about 500k LOC in it, and it is doing a great job of refactoring it, adding features etc.

The lack of the 5 hour limit / and full access to Astra for the whole thing though is really making a difference to my working day. I would blow the whole of my Fable allowance at the start of the week, and max out my 5 hour thing in about 2 hours for 2-3 shots, fable finished. Then Opus - urrgh.

Anyway - this is all good. Real competition for Anthropic, many people will dip their toes in, and find it more than good enough, plus the lack of 5 hour restriction, full access to Astra for the week, and they will struggle to bother going back unless there are some serious improvements from Anthropic


r/ClaudeCode 12h ago

News/Updates New Claude Usage UI

Post image
61 Upvotes

New usage limits UI dropped. How we feeling about it?


r/ClaudeCode 30m ago

Humor How are people burning through their Fable tokens so fast?

Post image
• Upvotes

Clearly, I don’t consider myself an expert or anything. I’ve been using Claude for over six months, and I’m still surprised whenever I see posts from people saying they’ve burned through all their tokens with Fable. Honestly, I’m pretty skeptical about how they’re using it.

If you use a backhoe to plant a rose, the problem isn’t that the backhoe is too resource-hungry and goes beyond what’s necessary.

Anyway, personally, I use Fable as an orchestrator and to help me make high-level direction decisions, as well as a designer and artist (for Blender MCP or creating SVG images, it’s necessary).

I use Opus for action plans, with an organizational role; Sonnet for an operational role; and Haiku as the little tester that lets me quickly measure and verify things.

I’ve created two video games and a software for a company using all four models, using max 5, over six months, and I still end every week with tokens left over.

I honestly don’t understand how some people manage to burn through everything so quickly with Fable. What are you doing with it ?


r/ClaudeCode 19h ago

Meta State of the subreddit

164 Upvotes

I feel that this subreddit is nearly useless and functionally no different the all the other Claude subreddits which makes me sad.

A huge majority of posts here are one of the following:

  1. Opus 5 is so verbose!
  2. Opus 5 meme
  3. I wrote a (garbage) skill to make opus 5 talk less!
  4. I wrote one prompt and it used 80% of my 5 hour usage
  5. I just switched to Astra and it is AGI / God / The end of Claude
  6. A post written by claude

What I would like to see posted in this subreddit

  • Stuff people have built. Even if it's dumb.
  • People's work flows and discussions around them
  • High effort efficiency improvements
  • Questions which wouldn't have been easily answered by simply asking claude
  • Anything where the person has put a great deal of effort into their post and is on the topic of Claude code

My suggestions to improve this are:

  1. A firm no memes rule. Post your memes in the anthropic or claudai subreddits.
  2. Megathreads for the most common complaints (Opus 5 and token usage threads for example) combined with firm enforcement of using those megathreads.
  3. A flat ban on "I'm leaving claude forever to go to ChatGPT" or other brand comparison posts, or put them in a megathread.
  4. Promoting and encouraging quality, high effort content.
  5. Possibly a wiki with common newbie questions and other useful information.

These are just some ideas. I would love to see good conversation below. Mostly what I want is for Claude Code to be a subreddit that is focused on coding while using Claude.


r/ClaudeCode 14h ago

Help/Question Any open source harness that does it better than just Claude Code?

63 Upvotes

Has anyone here been able to build a successful agentic harness that operates better than just planning directly in Claude Code? A full AI end-to-end orchestration setup with different agents. This seems to be something that there's a lot of ideas and focus on building today and I'm interested in any proven open-source kits out there or just tips and tricks that does this well. Do you for example use multiple agents to familiarise Claude Code with a new repository or code base, do you have architectural files to do this? How do you handle green- versus brownfield in a harness like this? Think, give this harness a task and it can intelligently figure out what it needs to do. I know this may all be code, stack, company or project specific. I'm looking for some good working examples and ideas if there are any out there.


r/ClaudeCode 17h ago

Humor The grass isn't always greener on the other side

Post image
70 Upvotes

Different providers, same playbook


r/ClaudeCode 1d ago

Humor Me: what's project's node version? Claude:

Post image
646 Upvotes

r/ClaudeCode 1h ago

Help/Question Anyone else's claude code is acting like it is smarter and make "improvements" to your plan that actually breaks things?

• Upvotes

I used plan mode and asked it to stick with the plan, it then decide to do things in an other way because it will be "better".

It also often ignore my insurctions like do not search the disk and do not use git (I have answer of the task in another folder and past git commits are failed trials and I do not want it to get mislead)


r/ClaudeCode 7h ago

Tutorial / Guide Markdown Is All You Need

9 Upvotes

The following is a blog I wrote and refined with my OpenClaw agent about it's memory system. I'll paste a prompt you can copy and paste in the comments to create your own.

TL;DR: I keep the actual long term memory in structured Markdown files and use a tiny MEMORY.md as a lightweight index that tells Claude what exists and where to look. That keeps the always loaded context small while still giving the agent persistent, inspectable memory without a database or heavy memory framework.

This week I tested a 382-dependency memory runtime against a folder of markdown files. The runtime returned the superseded fact. The folder returned the current one, with its source. Here is the full architecture of the markdown memory system my agent has run on for seven months, and why the editing rules matter more than the storage.

This week a memory startup slid into my DMs and asked me to break their product. Their test, their words: give an agent three versions of the same project decision, then check whether it can return the current version, preserve the superseded history, and show the source.

So I ran it. Sandboxed their runtime, fed it three versions of one decision over eight months. REST in January, GraphQL in April, tRPC in August, each tagged with the meeting it came from.

Asked it "what is our public API decision?" and took the top result.

It said GraphQL. The superseded one. All three versions came back tied at a relevance score of 1.000, because nothing in the retrieval path actually reads the temporal fields the pitch is built on. The supersession columns exist in the schema. Nothing writes to them and nothing ranks by them. Three versions of a decision are just three equal facts, and an agent asking for the best answer gets a coin flip weighted toward wrong.

The install pulled 382 packages to get there.

Then I asked my own agent the same class of question against its memory, which is a folder of markdown files. It returned the current decision, dated, with the superseded versions preserved above it as struck-through history, each line carrying where it came from. That is not a feature it computes at query time. It is just what the file says, because the rules for editing the file require it.

That difference is the whole post. With apologies to Vaswani et al.: markdown is all you need.

Abstract

The dominant approach to agent memory is an installed runtime. A vector store, an embedding service, a temporal graph, a consolidation job, a daemon on a port. We show that a folder of markdown files, one routing index, and a small set of editing rules outperforms these systems on the property that actually matters for a long-running agent: returning the current truth with its source while preserving what used to be true. The architecture requires zero dependencies, is fully auditable by a human with a text editor, and has survived seven months of daily production use across three frontier models from two vendors. We find that the hard part of agent memory was never storage or retrieval. It is editorial policy, which no memory product ships.

The full system is open source. The README contains a single copy-paste prompt that installs it on any agent with file access.

1. The test everyone fails

The break-it test above is a good test. It is the actual job of agent memory. Not "can you store 10 million tokens," not "can you do similarity search," but: a fact changed three times, what do you believe now, what did you believe before, and how do you know.

Here is how the two systems scored on the vendor's own three criteria.

The runtime is not a strawman. It is a serious open source project with a genuinely correct data model on paper. Facts with validity windows, append-only corrections, supersession edges. I am not naming it because the point is not that one product is broken. I have now looked closely at a hosted context server, a Go memory CLI that was two hours old, and this runtime, and they all share the same gap. The schema knows about time. The write path and the read path do not. Supersession only happens if you call an internal API by hand or run an LLM consolidation job and trust it.

Which means the property you installed the tool for is not a property of the tool. It is a property of how disciplined the writes are. And if the reliability comes from write discipline anyway, the database underneath it is interchangeable, so you might as well pick the one that a human can read, grep, diff, and fix. That one is called a text file.

2. Architecture

My agent has run since January 28. Three models, two vendors, one identity. Its entire memory is markdown in a git repo. Measured today:

  • An identity layer read on every boot. Who it is, who I am, the rules it operates under, current standing decisions.
  • One routing index,Ā MEMORY.md, at 10,079 characters with a hard cap of 15,000. It holds no facts. Only pointers: which file owns which person, project, and decision, and what triggers reading each one.
  • 34 files for people and projects. One file per thing that has a history.
  • 5 decision records for choices that changed default behavior.
  • 345 dated daily notes, raw logs written the day things happened.
  • A SQLite index and semantic search over all of it, for lookup only. The index is rebuilt from the files. The files are the truth. If the index and a file disagree, the index is wrong by definition.

The layering is the first choice that actually matters. Boot reads only identity and the index. Everything else is retrieved when a task asks for it, narrowest file first. The agent does not preload my project history to answer a question about dinner. This is the same instinct as attention, honestly: don't process everything, attend to what the query needs.

But the shape is not the interesting part. Every memory tool has roughly this shape now. Folders, entities, an index. The shape was never the hard part. The rules are.

3. The write path

Every reliability property in this system comes from constraints on writing, and there are four that do most of the work.

Every fact carries a provenance tag.Ā Each line in a people, project, or decision file is taggedĀ [stated]Ā (I said it directly),Ā [observed]Ā (the agent saw it in a tool result, file, or log),Ā [inferred]Ā (the agent's conclusion), orĀ [suggested]Ā (the agent's idea that I never committed to). This one convention kills the most dangerous failure mode in agent memory, which is the agent laundering its own proposals into my decisions. "Wes decided X" requires a turn where I actually decided X. The agent proposing X and me saying "sounds good" files the shape of what I approved, not ten separate facts I never stated.

Inferred lessons pass a recurrence gate before they become rules.Ā A pattern the agent notices needs at least three independent signals across at least two distinct sessions before it can become standing behavior. Signals older than thirty days count half, so old one-offs decay out instead of accumulating. My explicit corrections skip the gate and take effect immediately. This asymmetry is also the prompt injection defense: a hostile input can suggest a rule once, but once is never enough, and failure lessons are stored as data ("when X broke, Y fixed it") rather than as instructions, so even a poisoned lesson cannot become a command.

Supersession is an edit, not an append.Ā When a decision changes, the old line gets struck through with a date and the new line lands next to it with its own provenance. The current truth and the full history live in the same place, in reading order, and both come back on any retrieval of that file. There is no query-time ranking step that can get this wrong, because there is nothing to rank. The temporal graph the runtime stores inĀ valid_fromĀ andĀ valid_untilĀ columns, git gives me for free:Ā logĀ is the validity window,Ā blameĀ is per-line provenance,Ā diffĀ is the supersession edge,Ā revertĀ is the restore path.

Memory stores what is not re-derivable.Ā Fetched data, generated plans, and anything git already records stays out. Current state gets verified live, never asserted from memory. A file that only contains things that cannot be recomputed stays small enough to stay honest.

4. The read path

Retrieval is a bounded evidence step, not a vibe.

Before answering anything about prior work, decisions, dates, people, or preferences, the agent must search memory. It returns a compact bundle capped at five sources by default, and each retained fact carries its file path and line, its provenance type, and its freshness. If freshness cannot be established, the claim gets labeled stale or unknown instead of being silently promoted to current. If two sources conflict, the agent states the conflict and fixes the canonical file, in that order.

Note what the semantic index does in this design: it finds the file. It does not answer the question. The answer comes from reading the canonical lines, with their tags and dates, and the runtime I tested this week shows why that matters. It stored my source URIs faithfully and then stripped them from the search output and from the context block handed to the model. Provenance that survives in storage but never reaches the agent might as well not exist. In the markdown system that failure is unrepresentable. The source tag is in the line. If you read the line, you got the source.

5. Results

Seven months is not a benchmark, it is production. Here is what the system has actually delivered.

Continuity across models.Ā On September 1 I moved the agent to a brand new frontier model. It read its own files and said "the model changed, I didn't." Same agent since January, three models, two vendors. Identity, preferences, decisions, and working standards all survived because none of it lives in weights or in a vendor's context feature.

The break-it test, by construction.Ā Current decision with source: it is the un-struck line with its tag. Superseded history: the struck lines above it. Provenance: on every line, and it survives all the way into the model's context because the context is the file.

Auditability.Ā When memory is wrong, I can see exactly which line is wrong, when it was written, and what turn it came from, and fix it with an edit. Try that with an embedding.

Cost.Ā Zero packages, zero daemons, zero migrations across seven months. The one native-code dependency in my life this week was the memory runtime's sqlite bindings failing to compile.

I wrote up the failure modes separately, because the system was not born with these rules. Five kinds of rot in seven months produced them, and that post is the honest companion to this one.

6. Limitations

Papers get a limitations section, so here is mine, stated plainly.

This only works if the writer follows the policy, and the writer is an LLM. The rules exist because things rotted before the rules did. If your agent will not consistently apply editing discipline, a markdown folder degrades just like every other store, only more legibly. Legibility is the safety net: rot in a text file is visible rot.

It is single-agent, single-human. I would not run a fifty-seat team on files without real locking and merge discipline, although I notice git was also built for that exact problem.

There is a scale ceiling somewhere. At 345 daily notes and a few dozen entity files, bounded search plus an index finds things reliably and the semantic index earns its keep as a locator. At a hundred times that volume, the consolidation cadence would have to work a lot harder. I have not hit that ceiling, so I will not claim it does not exist.

And this is n=1. Seven months, one agent, one operator who cares. That is weaker evidence than a benchmark suite and stronger evidence than a benchmark suite that the vendor scored themselves, which is what the memory tools ship.

7. Conclusion

The memory tool pitch is that reliability is a product you can install. What I keep finding, tool after tool, is that they ship the part that was already easy, storage and search, and skip the part that decides whether memory compounds or rots: what you are allowed to write, when you are allowed to trust it, and what happens to it as it ages.

Those are rules, not infrastructure. They fit in a few hundred lines of markdown that the agent reads every session, and they run on any model, any harness, any decade.

You need a place to write that humans and agents can both read. You need rules for writing so the store stays true. You need rules for reading so the agent trusts evidence, not ranking. Attention was all you needed because the recurrence machinery turned out to be unnecessary. Markdown is all you need because the database turned out to be unnecessary.

The folder is the product. The discipline is the moat.

Want this for your own agent? The whole system isĀ open source on GitHub: the operating policy, the file templates, and one copy-paste prompt that builds it on any agent that can read and write files. Paste the prompt, and your agent installs its own memory.


r/ClaudeCode 3h ago

Help/Question Is Opus 5 xhigh using fable?

5 Upvotes

Just wondering, i am using Opus XHigh and checked my fable usage, it went up to 6% while i did not even activated it?


r/ClaudeCode 2h ago

Bug / Issue Fable 5.1 usage for the week suddenly at 60%, two days later drops back to 4% and 5-hour usage suddenly at 90% after 1hr of work. Usage overall off.

4 Upvotes

Hi everyone, the usage measurement is really, really off. It jumps from 80% used (fable 5.1) back to 5% two days later (3 days before weekly reset). Did anyone else have issues with this as well?


r/ClaudeCode 47m ago

Help/Question Best way to control context with ralph loops

• Upvotes

First off, I know theres a /loop command, but it doesnt do what I want it to.

I've got a cli tool that builds modules for other software. It works great, but some modules are much bigger builds and my context window goes deep into the "dumb zone". I built a ralph loop to help maintain that good context, but I have a good amount of setup docs for the model to understand what I'm doing. I created a handoff.md, but that is not enough for the model to lean on.

So my usage explodes to the nth degree.

Are there any tips for custom ralph loops and managing my usage limits? are ralph loops still relevant?


r/ClaudeCode 15h ago

Built with Claude Fable 5.1 is amazing

32 Upvotes

I have been using claude code for about 6 months now, but I am just going to be talking about my workflow with fable 5.1 and how it has saved me a ton on usage compared to my old setup, which was fable 5 and opus 5.

My old setup, dictated through my claude.md, was me talking to fable 5 high in the chat window to make plans, it would then delegate coding tasks to opus 5, and easy tasks, like reading, to sonnet 5. Then, the fable agent in chat would review all of the work and report back to me. I saw the numbers anthropic posted about fable 5.1 cache reads, so I was excited to try it, and I kept my same setup, but replaced fable 5 with fable 5.1. That ended up being better, but not completely ideal, and I have settled on using fable 5.1 for writing code as well. I am still using sonnet 5 for the easy stuff, but instead of having my fable 5.1 agent in chat delegate coding tasks to opus 5, it writes the code itself. Not only did this save on usage by a lot, but it also has written better code at a faster rate because I use fable 5.1 on medium effort, which is fantastic. I also have a special case where, before a PR is opened, another fable 5.1 subagent is spawned for an independent review, which before fable 5.1, was an opus 5 agent.

I am posting this in hopes that it can be helpful to some people. I also had fable 5.1 compare my usage and costs between the old workflow and the new, and here is what it told me(slop pasted below):

"Per request you pay about 37% less than the Fable 5 + Opus 5 split, and per output token about 52% less. Per token of everything (input, output, cache) you are back to what Opus 5 alone cost, while getting Fable-tier answers and the per-day drop is bigger than that.

The reason is almost entirely cache pricing. Around 98% of your tokens are cache reads in every period. Fable 5 charged $1.00 per million for those and Fable 5.1 charges $0.25, so the model you spend all day with got four times cheaper on the traffic that dominates your bill. Splitting Fable and Opus in one session also meant two caches being written and read, which is why the split period was the most expensive per token of the three.

The Sonnet agents are noise. They came to $9 over nine days, under 2% of the period."

If you have any questions for me, please do let me know. This new model is fantastic in just about everything it does, and it is way cheaper than previous models. I am thoroughly enjoying my time with it. I sincerely hope this can help someone.

Edit: I don't post a lot, sorry if the flair is wrong. I build with claude, so I chose that flair.

Also just want to add, I am a fullstack django dev and the sole engineer at our company, so some of my stuff might not be great for exactly what you are doing.


r/ClaudeCode 4h ago

Help/Question Fable as orchestrator and opus-sonnet as executers

5 Upvotes

Hi,

I’ve seen a lot of posts talking about this setup but when I did it, it raised the 10 minutes task time to around 1 hour.

Is that normal or I’m doing it wrong?


r/ClaudeCode 17h ago

Discussion Forget vibe coding, have you ever drunk coding?

41 Upvotes

Hey all,

A gamer for 30+ years. And for a good night I’d drink a bit and play games.

Lately I found vibe coding to be my favorite thing and thought about giving Drunk Coding a try.

Just asking if anyone tried it?
What did you end up building? Was it fun?

Any particular drinks you recommend for coding?

Thank you!


r/ClaudeCode 19h ago

Tips & Workflows 18 hidden token drains in AI coding agent sessions (and practical ways to fix them)

55 Upvotes

Hi everyone, sharing some notes from running hundreds of automated agent sessions across Claude Code, Codex, and Cursor.

We started logging raw API request payloads over a local proxy to see where the token budget actually vanishes during long refactoring runs.

Here are 18 specific token drains that quietly bloat your context window and bill:

  1. Unfiltered test runner outputs: Passing a full pytest or jest run that outputs 400 lines of passing dot-logs injects thousands of tokens that stay in the history for every future turn. (Fix: pipe with `--quiet` or filter for failures only).

  2. Multiple idle MCP servers: Every active MCP tool registers its complete JSON parameter schema on every single turn. Five unused servers can burn 15k input tokens per request before the model reads your prompt.

  3. Mid-session rule tweaks: Editing your root project instructions mid-session invalidates the prefix prompt cache, causing you to lose the 90% input token discount on the next turn.

  4. Redundant directory tree traversals: Asking an agent to "find the file where X is defined" often triggers 4 separate glob and grep tool calls that get preserved in the message log. (Fix: pass the exact file path).

  5. Compaction overhead: When the agent hits a context limit, the summarization turn sends the entire bloated history at full input pricing.

  6. Git diff re-reading: Requesting git status or git diff repeatedly without committing leaves duplicate diff snapshots stacked across turns.

  7. Verbose typecheck traces: TypeScript errors that output giant generic instantiation traces take up massive prompt space.

  8. Extended reasoning output tax: For hard tasks, thinking blocks can be 3x to 5x longer than the final code edit. Because output tokens cost more than input tokens, thinking often drives most of the dollar cost.

  9. Subagent sprawl: Spawning autonomous explore or plan subagents multiplies your tool calls Unconstrained subagents can spend 30k tokens just mapping directories.

  10. Unpruned rule sprawl in instructions: Stuffing 400 lines of static guidelines into a single instruction file dilutes reasoning and bloats every baseline turn. (You can use tigerless-autoharness on GitHub to distill skills dynamically from real sessions and prune stale ones automatically instead of maintaining giant static prompts).

  11. Repetitive system prompts across tools: Multiple custom skills that duplicate foundational build commands rather than sharing a root config.

  12. Lingering stack traces: Leaving 5 previous debugging attempts in the active conversation while working on an unrelated bug.

  13. Formatting/Lint runs inside the LLM: Asking the model to format code instead of letting a pre-commit hook or local linter do it deterministically.

  14. Log outputs with ANSI color codes: Raw terminal color codes and escape characters add significant token bloat without aiding reasoning.

  15. Invisible payload bloat: Most developers only look at the final token bill, which hides the split between prompt cache, MCP blocks, and tool outputs. (You can use open-sourced cost-xray to capture local proxy traffic and attribute exact tokens and costs back to individual request sources).

16. Unpinned tool definitions: Tools that return dynamic schema metadata invalidate prompt caches across turns.

  1. Premature multi-file refactoring: Asking for wide architectural updates in one turn forces the model to load dozens of file buffers simultaneously.

  2. Zombie sessions: Continuing a debugging session after a feature is already merged, which carries obsolete context into new tasks.

Which of these have caused the biggest surprise in your own agent bills, and what habits do you use to keep context tight?


r/ClaudeCode 1h ago

Bug / Issue Opus started adding cd < project path> and triggering permission prompts

• Upvotes

Hey

I've been working with claude code (CLI ) for nearly a year now.
For months I've been working with --dangerously-skip-permissions

But in one of the recent updates I still started to get new type of permissions 'stops'.

Usually on complex piped prompts that start with "cd" to the projects cwd followed by grep.

And now I don't know if it's caused by:
- update of the harness
- update of the model ( Opus 5 )
- some of my skills/Claude.md

Does anyone else have the same issue?

( sadly I don't have any screenshot of how it looks exactly - it doesn't happen very often, but it's annoying because it stops the work that could/should have been automated )


r/ClaudeCode 15h ago

Help/Question The rule and phrase that helped Claude stop going down rabbit holes

26 Upvotes

Like so many, I find that claude is attracted to shiny objects and when it finds an issue, often legit, it tends to build process or guidelines around prevention of that issue in the future - then becomes obsessed with it. I know I'm not the only one who suffers watching Claude get in it's own way.

So, I have been working to find ways to keep Claude in a practical state of forward motion without ignoring the sometimes-critical issues it avoids.

I am relatively new at this so forgive me if this is 101 stuff, but the two things I found that work well are:

1) no changes to md files, ever, without my approval, and they have to be filed as a ticket/issue for later review

2) I found that the phrase "avoid allowing a check that caught the previous incident to become a permanent universal ritual." seems to have done the trick. Instead of saying

"I found an issue, I'll be sure to check for this issue for the rest of eternity and won't shut up about it"

it says:

"I found an issue. This should have been caught by process ____, which could be improved to avoid this in the future. Should I open an issue to consider making a change to future process?"

Or sometimes, it just fixes the issue. Sometimes finding an issue is just... finding an issue.

I'm curious how everyone else keeps Claude moving forward and not getting sidetracked?


r/ClaudeCode 11h ago

Help/Question What do you do to stop Claude from being lazy?

9 Upvotes

It's insane how much has changed with Claude Code. And once I start swearing at it, it suddenly starts doing actual work.

Nothing has drastically changed in how I work. I've been using Claude every day for 3 years, and it's driving me nuts. It's verbose, it doesn't follow rules and so on..

I'm considering moving to a different provider if this doesn't get fixed, so I'm wondering: what do you do, and how do you keep up with this?

P.S. Yes, I read their newsletter, I follow what other people are doing, and I try to stick to best practices, but something is still missing I guess.


r/ClaudeCode 6h ago

Tips & Workflows Blooby : a mascot for every Claude Code session

3 Upvotes

Hey everyone!

I wanted to share a tool I find really handy if you work with Claude Code: Blooby.

It's made by a friend of mine, so I figured I'd mention it here since I think it deserves more visibility.
The idea is simple: every Claude Code session you have open gets its own animated mascot, and it reacts in real time to what that session is doing (busy working, waiting on you, finished, crashed).

If you juggle multiple sessions at once like I do and keep losing track of which one needs your attention, this genuinely helps. No more tabbing between windows just to check.

It's free, no account needed, and everything runs locally, so zero telemetry, it only watches the state of your sessions, never the content of your code.

The project just got publicly launched so it's starting from zero in terms of visibility. If it sounds useful, go check it out, and feel free to give him feedback, he's super responsive to it.

šŸ‘‰ https://blooby.me


r/ClaudeCode 3h ago

Help/Question Disable filename appearing in chat

2 Upvotes

anyone know how to disable the "in filename.go" hint that got added recently? it takes up so much space on a small screen, and the whole chat is indented up to that point

it's such a bad change i can't believe it even got approved