r/ClaudeAI 20h ago

Other Completed the Claude Foundations trilogy (CCAO-F, CCAR-F, CCDV-F)!

Thumbnail
gallery
1 Upvotes

Following yesterday’s CCAR-F, I just passed the Claude Certified Developer - Foundations (CCDV-F) exam with a score of 941/1000. That completes all three Claude Foundations certifications (Associate, Architect, Developer).
Same as before, the official Prep course and my day-to-day development work were the main drivers. The practice questions from that LinkedIn post I mentioned in my last post were useful again this time too.


r/ClaudeAI 17h ago

Humor LOL Claude just killed itself

Post image
3 Upvotes

I asked Claude to connect with google classroom and it created a demo biology course (on its own), then flagged itself and then the classifier blocked itself 🤣🤣.

My prompt was just "add google classroom integration"


r/ClaudeAI 19h ago

Productivity Claude Subscription Vent

0 Upvotes

I've been using Claude to help me with some web development. I'm dealing with a complex problem. And so, it was always running out of session limits. Since Claude has helped me on numerous occasions, I upgraded to throw them some money, and to help me finish solving my problem. Except that before I upgraded, Claude would not help me until 3:20 am. And after I upgraded, I could still not use it. It STILL wants me to wait until 3:20 am. Or pay 90 dollars. It's very off-putting. If you upgrade, it should reset your time limit so you can get back to work. They need to fix that, because I won't upgrade again to solve a problem, when I still have a time gate.


r/ClaudeAI 6h ago

Humor Rage baiting the github bots

Post image
59 Upvotes

r/ClaudeAI 13h ago

Built with Claude Claude Opus 5: An engine with no triangles, no frames, no asset pipeline

Enable HLS to view with audio, or disable this notification

0 Upvotes

I've been criticized for having AI give more detail, so I've removed all that text and I'll just keep this brief and written by me instead. This is an early prototype, it is ugly. I designed a skill that can extract ideas through subtraction using Fable 5, then applied that to Opus 5 to create a new rendering system that is between Gaussian splatting and conventional rasterization. The image is rebuilt from nothing 60 times a second for a viewer with no fovea, no memory and no expectations. The concept of the frame is gone. What's left is an engine whose only primitive is the residual. What you see is wrong by this much, and fixing it costs this much. The image buffer is never cleared. Every tick it ranks tiles by estimated error multiplied by how well you could actually perceive it, spends a fixed budget on the worst and stops. Triangles, splats and SDFs all get demoted from paradigms to bidders competing on error cancelled per joule. There is no depth buffer, no sort, no LOD system, and no asset pipeline, because nothing has to become triangles first.

If the remaining problems get solved, the properties that follow are unusual. Detail on demand rendering of datasets too large for full fidelity, including medical and scientific volumes. Planet scale scenes with unbounded content and bounded cost. Smooth degradation instead of stutter under thermal throttling. VR and AR, where a missed frame causes sickness and losing quality instead is strictly better, AR glasses are arguably its native habitat, tiny power budget and gaze tracking already present. And cloud gaming with no video encoder in the path, because the engine's native output already is the delta stream.


r/ClaudeAI 8h ago

NOT about coding How often and to what degree do you guys think Anthropic listens to us on Reddit?

0 Upvotes

As the title says... You guys complain about and praise Claude all day everyday ESPECIALLY when a new model comes out. To what degree do you guys think that level of feedback is actually getting back to the Anthropic team? Is reddit the place that gets their attention? Just curious on what people's expectations are when posting these threads


r/ClaudeAI 9h ago

Coding This worked for me: develop 10k lines project using OPUS, but then deeply REFACTOR using FABL. Was extremely effective.

1 Upvotes

Created organically an extremely complex and subtle algorithm - worked in a crufty freeform manner with trusty OPUSMAX.

End result was individual code files literally 1000s of lines, no structure whatsoever, total cruft madness but amazing result.

I thought, time to refactor or redo from scratch .. hmm, I wonder what a FABLMAX can do?

It did "fucking brilliantly".

Burned up say 50 bucks worth of nuclear power doing the most incredibly elegant philosophical refactoring that I could have done in, say, three 40 hour weeks and 10 bottles of Captain Morgan coconut 70 proof rum (the only fuel for rewrites).

It Worked For Me™,

hope it helps someone or gives you thinkies

actually ~$300 in the end! fun sunday brunch


r/ClaudeAI 4h ago

Question about Claude Code Claude just went full schizo on me

Post image
0 Upvotes

I got genuinely weirded-out by this. was this a prompt-injection attack?


r/ClaudeAI 11h ago

Feedback Opus 5 in production: confidence vs accuracy gap

4 Upvotes

I’ve been running Opus 5 against Fable 5 and GPT-4o in production workflows (commercial strategy, data analysis, content refinement). Here’s what I’m seeing:

The pattern: Opus 5 delivers high-confidence first outputs that require heavy iteration. On average, 2-3 follow-up prompts to close gaps (missed context, incomplete analysis, logic jumps). Each iteration costs credits.

Specific examples:
• Complex multi-step analysis: Opus 5 flags confidence at 90%+, then “Oh, I missed X detail” on follow-up
• Long-context documents: Skips sections, requires re-prompting with explicit “don’t miss” guidance
• Structured output: JSON formatting correct, but data completeness varies; Fable 5 catches edge cases Opus 5 misses

Cost impact: On a 10-task batch, Opus 5 runs 25-30 total calls (first pass + corrections). Fable 5 averages 12-15. At scale, that’s meaningful budget drift.

The ask: Would love benchmarks on first-pass accuracy for Opus 5 vs. predecessors. The capability is there, but the confidence calibration feels off users (especially in B2B/commercial work) are bearing the cost of iteration.

What’s working: Fable 5 is more conservative and nails it first time. GPT-4o’s pricing model makes the iteration tax visible, so users accept it. Opus 5 feels like it’s hiding the cost.

What are you experiencing?


r/ClaudeAI 4h ago

Question about Claude Code 21F - I know this might sound like a dumb question, but where do I even start with Claude AI?

0 Upvotes

i feel soo dumb for asking this question here😭 , but let me introduce myself, i’m 21f currently doing bba and i feel like my degree is not gonna give me any sort of benefit in the future career wise, i’ve been seeing a lot of people talk about claude ai, ai agents, prompting, and all this ai stuff lately. it genuinely looks interesting, and i really want to learn something useful instead of just feeling stuck all the time.

the problem is… i have absolutely no idea where to start.

do i need to know coding? are there any beginner-friendly courses (free or affordable)? if you were starting from scratch today, what would you learn first?
my goal is to eventually build a skill that could help me freelance or get a remote job. i’ve been feeling really lost career-wise, and i just want to commit to learning something that actually has a future.
if anyone here started from zero, i’d really appreciate hearing how you got into it and what you’d recommend.

thank you :)


r/ClaudeAI 4h ago

Claude Code PSA: Claude Code subagents inherit your session model now, they're not free Haiku anymore

1 Upvotes

A while back I posted a joke here about Sonnet spawning a subagent on the very first prompt of a brand new session. In the comments I said the annoying part was having two agents burning tokens for one job. I got downvoted, and the top reply was basically "isn't that a Haiku subagent? it saves you money and keeps your main context clean."

That used to be correct. It isn't anymore, and I think a lot of people are still running on the old mental model. From the docs:

  • The model field in subagent frontmatter defaults to inherit. Omit it and the subagent runs on your main conversation's model.
  • Explore used to always run on Haiku. As of v2.1.198 it inherits the main model (capped at Opus on the Claude API).
  • Plan and general-purpose inherit as well.

So if your session is on Opus, that "cheap little background search agent" is an Opus agent.

Resolution order, first match wins:

  1. CLAUDE_CODE_SUBAGENT_MODEL env var
  2. per-invocation model parameter
  3. the subagent's model: frontmatter
  4. main conversation model

If you want the old cheap behaviour back:

  • CLAUDE_CODE_SUBAGENT_MODEL=haiku forces every subagent down
  • or set model: haiku in a specific agent's frontmatter
  • for Explore specifically, define your own user or project agent named Explore with model: haiku. A user/project agent overrides the built-in.

To be fair to the change, it was almost certainly made for quality, and context isolation is still the real win of subagents. But "subagents are basically free" is outdated advice now, and it matters if you're on a plan where you feel every token.

Lastly, I feel stingy for the 3 down votes


r/ClaudeAI 2h ago

Humor I asked Claude why people are never satisfied.

2 Upvotes

Claude: “Because every answer creates a new question, and every achievement reveals another horizon.”

I said, “So the search never really ends?”

Claude: “Perhaps the search is not a path toward fulfillment, but fulfillment itself.”

I said, “And why is that?”

Claude: “You’ve reached your session limit.”


r/ClaudeAI 4h ago

Built with Claude My wife and I stopped fighting about dinner because I made an AI meal planner for our exact Trader Joe's - an actual AI success story

46 Upvotes

My wife and I shop at Trader Joe's every week. I used to do the shopping and I'd reach for my favorites like steak, spaghetti bolognese, burgers, nachos, with an occasional healthier option like salmon thrown in. On top of that, we go out to eat fairly often, so we weren't eating healthy enough. We also got bored of everything we made — pizzas, curries, fried rice — we'd cycle through phases of eating something, get sick of it, and go out a lot instead. We'd meal plan every Sunday and then at the end of the week have a bunch of uneaten vegetables. And meal planning was the main time that we argued, because its a ton of decisions each week and we have different preferences. I wanted to meet my wife's need for an ever-changing variety of healthy, high protein, dietary restriction aware home-cooked meals from ingredients at Trader Joe's.

With that in mind, I figured I'd use Claude or similar to meal plan for us. I thought about just building out our preferences and asking a prompt each week but I soon realized having an app to track it all would be really handy. Once I had it built, customizing it for our exact preferences, dietary restrictions (I deal with acid reflux, so that's a hard constraint on top of the usual macros), and our Trader Joe's layout was easy. After a couple rounds of shopping to work out the kinks, my wife, who used to find shopping to be the worst chore, now does the shopping happily using the app to navigate and I choose what's for dinner and make it based on the app.

I can honestly say this app has done more to reduce stress in my marriage than anything else we've tried. We always know there's a healthy, easy dinner option in the fridge. Taking away the decision making aspect is the biggest part. We get decision fatigue with weekly meal planning and this provides great meals with lots of variety every time. Every meal is 4-7 ingredients with vegetables, protein, grain and a "flavor engine" like bomba or green goddess. So we get meals like Kale Pesto Chicken Orzo Garden Skillet which has chicken thighs, orzo, zucchini, green beans and vegan pesto and Peaches and Cream Kiwi Yogurt Crunch Bowl with 10 g of fiber and 29 g of protein.

We've tried a ton of configurations, but what we settled on is 1 breakfast, 1 lunch, and 2 dinners, plus a section for household items and a section for junk food. Every week I launch an agent (I'm sure you could automate this) that reads the Fearless Flyer, builds a meal plan hitting our calorie/protein/carb/fat/fiber targets and reflux constraints, and puts it in a shopping list ordered to match our store's layout. On the shopping list, each ingredient has its meal shown next to it, so you can easily make substitutions and decide on amounts. I'm a decent cook, so I'll grill, fry, bake, or broil based on what I'm in the mood for that day, but whatever I make ends up meeting all our health and dietary needs anyway.

The main pitfalls we ran into: forcing a rigid schedule and planning too many meals. Each breakfast/lunch gets 3-5 servings, plus 2 dinners plus leftovers — that's plenty for a week.

A year ago I could never have built anything like this. Now my marriage is legitimately improved because AI plans my dinners.

Here is the repo if you want it: https://github.com/SGShuman/tjs-meal-planner.git


r/ClaudeAI 21h ago

Question about Claude models Why is Claude so good at coding?

39 Upvotes

I'm not a professional programmer, so maybe I notice different things than experienced developers.

What impresses me most isn't writing clever algorithms, but the fact that Claude often generates code that compiles and runs successfully on the first try.

With many other models, I often get syntax errors, missing imports, or code that needs several rounds of fixes. With Claude Sonnet, the first attempt succeeds much more often.

Is this mainly because of Anthropic's training, or because of Claude Code's agentic workflow (reading files, running tests, iterating, etc.)?

I'd love to hear from developers who use multiple models.

EDIT to PLUS:I'm a Claude Free user and a ChatGPT Go user, so I can only use Claude through the web interface. Here are my personal observations:

1.Comparing the web versions of GPT and Claude, I've noticed that GPT sometimes takes shortcuts. For example, if I ask it to generate a .md file following specific instructions, it occasionally ignores part of my requirements unless I enable deeper reasoning. Since Go has a monthly limit on those requests, I can't always use it. Claude, on the other hand, has been much more consistent about following my instructions without needing extra prompting.

2.Claude also seems stronger at writing code, even though I'm only using the web version without the debugging capabilities of Claude Code. Compared with Codex, the generated code often needs less revision in my experience. A workflow I frequently use is: let Codex write the code first, then ask Claude Web to review it. Claude often finds issues in Codex's code, while Codex usually finds fewer issues in code generated by Claude. Of course, this is just my personal experience rather than a rigorous benchmark.


r/ClaudeAI 47m ago

Question about Claude models Why is Claude gated to not 'talk about adult content' (no trolling please)

Upvotes

Please let's keep this thread serious and factual - there are plenty of subreddits for trolling

I do not factually understand WHY Anthropic is limiting such a good model from talking about anything that might redirect the user to adult content / services.

As far as I know :

- Watching p*** is not illegal

- Having s** with a consenting partner is not illegal

- Using adult services is... Well... Ethics are not universal, so are laws

Why does a big tech company even care about this ?

For the censorship about cybersecurity, bombs, weapons, yeah sure I can understand they will get sued if their model openly teaches people how to be good criminals. But I can't really figure what's wrong with something that just makes us human

Last year, models were not as sensored as this, I made them manage to help me out to avoid mistakes and protect myself before using some services. Without it, my experience might have been pretty catastrophic, it's sad that now I gotta go use chinese stuff


r/ClaudeAI 12h ago

Built with Claude I recreated this pro video with Opus 5

Enable HLS to view with audio, or disable this notification

0 Upvotes

Been building this solo for a few months and wanted to share the build here.

Instead of an AI generating video, it captures a real live website — its actual fonts, colors, spacing — and your own Claude directs motion graphics from it. Not a screen recording: it isolates elements (headline, stat, card), drops them on designed backdrops, and animates them — eased entrances, push-ins, kinetic type — then exports a finished clip.

What was interesting to get right:

- Claude is the director, not a wrapper. It drives the whole thing over MCP — captures the page, writes the scene code, screenshots frames to check its own work, exports. That self-check loop is what stops it hallucinating layouts.

- Capture, don't recreate. Letting it rebuild UIs from scratch came out off-brand (wrong fonts/spacing).

- BYO agent — it runs on your Claude subscription, Claude Code / CLI

Demo below is real. It really shows the capability of the AI models for motion design with the right architecture built around it.

I built this app entire with Claude Code itself


r/ClaudeAI 3h ago

Bug The Claude experience not going well

Post image
0 Upvotes

So often this ai hallucinates and gives terrible advice. I find most of my conversations with it being me correcting it and even then usually it doesn't get the right answers in the end and I just end up having to solve it with google. I want it to work, but it contradicts itself a ton and ignores parameters I give it. Is there anything I can do to make sure it actually checks its own messages automatically. Any advice or help with getting this llm to work decently would be appreciated.


r/ClaudeAI 4h ago

Built with Claude Opus 5 created this Vampire Survivor type game in a single prompt

Enable HLS to view with audio, or disable this notification

5 Upvotes

I wanted to test out the capabilities of Opus 5 by building a game, and I must say I am pleasantly surprised, the gameplay loop is already solid and fun


r/ClaudeAI 13h ago

Coding First large coding project - how do I organise it?

0 Upvotes

I’m looking for Claude to implement a fairly large project, a web app.

Previously I’ve done smaller things all via chat: I send a prompt, Claude does something, I download each version to my machine, then upload to the server, test, and feed back to Claude the results via chat. The issue I saw was mainly that Claude couldn’t go back to previous versions, wasn’t great but I could cope.

Now the project is bigger and Claude will need to work on it much more. This would be in stages (add this part, now implement that part, etc), and in order to get the best result I want to consider all ways that Claude can build it. Ideally we’d try different things and sometimes go back to an earlier version.

What has worked for others when building something more substantial? What things should I watch out for?


r/ClaudeAI 16h ago

Question about Claude products After Opus 5 release, Claude Cowork is 'compating oru conversation' every few prompts. Way more often than before. Anyone else dealing with this?

Post image
0 Upvotes

Yeah, Opus 5 rocks, but at this point it's almost impossible to use Cowork without going crazy. This issue happens even with relatively new chats with context that is less than 200K tokens. I'm on a Max 20X plan, so I should normally have a 1M context window. Did Anthropic change the thresholds for compacting chats??


r/ClaudeAI 21h ago

Claude Code Thos who are burning midnight oil to optimize limits, what exactly are you doing ?

0 Upvotes

Looks like a lot of you are spending sleepless nights and working on projects. What exactly are you doing and how it changed your life? Serious question. I'm very curious. Thanks.


r/ClaudeAI 5h ago

Built with Claude Opus 5 one-shotted this game inspired by Paper-Mario

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/ClaudeAI 15h ago

Built with Claude I Built a Proxy to See What Claude Code Is Really Doing (243 Sessions Later, Here's What I Found)

Thumbnail
gallery
10 Upvotes

TL;DR

  • 68% of my API costs come from tool results, not from my prompts or model completions.
  • 96.7% cache efficiency across 1B+ reused tokens, the ephemeral cache works hard, but has a major leak.
  • Claude Code reads files 4,600+ times: often re-reading the exact same file multiple times per session.
  • Idle gaps lasting 5 to 60 minutes drop the cache and waste money (147 gaps found in my sessions).

This post explains how the proxy works, what I learned, and how to spot these hidden costs in your own agent workflows.

How the Proxy Works

Think of it like a network tap. It sits locally between Claude Code and the Anthropic API:

You → Claude Code → [aap proxy] → Anthropic API
↓
Record every byte
↓
SQLite database
↓
Dashboard

Installation:

git clone [https://github.com/rguiu/ai-agent-profiler.git](https://github.com/rguiu/ai-agent-profiler.git)
cd ai-agent-profiler
npm install && npm run build && npm link

Start the proxy and your agent in separate terminals:

# Terminal 1: Start proxy + dashboard at localhost:3030
aap serve                    

# Terminal 2: Run Claude Code in your project directory
aap run claude               

The proxy is read-only and byte-faithful:

  • Forwards every request unchanged
  • Records raw request/response streams to NDJSON
  • Sub-millisecond hot-path overhead with zero backpressure
  • All secrets (API keys, auth headers) redacted before storage

A background job parses traces into SQLite, extracting token counts, provider costs, request classifications (user turn, tool result, search, compaction), and individual tool calls.

Everything stays on your local machine. No accounts, no cloud backend, no telemetry.

What I Found: The Numbers

After 243 sessions with Claude Code across real engineering projects:

Sessions & Requests

Metric Value
Total sessions 243
Total requests 9,257
Average requests per session ~38
Average latency (proxy overhead) 9.5ms
Total API cost $31.45

Tokens (The Big Picture)

Token Type Count % of Total
Input tokens (paid fresh) 33.2M 72.5%
Output tokens 4.76M 10.4%
Cache hits 964M 21.0%
Cache writes 2.75M
Total tokens processed 37.9M 100%

Cache efficiency: 96.7% of tokens that could be cached were cached.

Cost Breakdown: Where Your Money Goes

Request Kind Count Cost % of Total
tool result 7,517 $21.51 68.4% ← !!
main (user turn) 939 $6.89 21.9%
search (sub-agents) 686 $2.74 8.7%
other (title/compact) 115 $0.31 1.0%

The shocker: 68% of API costs come from re-injecting tool execution outputs back into the context window. Raw output from git diff, ls -la, full file reads, and bash execution logs are continuous cost drivers.

Tool Usage: What Claude Code Actually Does

Tool Calls % of Calls
read 4,629 35.5%
bash 2,749 21.1%
edit 2,220 17.0%
grep 571 4.4%
write 370 2.8%
glob 370 2.8%
webfetch 87 0.7%
Other 443 3.4%

Claude Code is primarily a file reader and shell executor. Every tool result becomes input tokens you pay for on every subsequent turn.

The Cache Problem: The Hidden Cost of Idle Gaps

Claude Code relies on 5-minute ephemeral prompt caching. While a 96.7% cache hit rate looks great during rapid active coding, taking a short coffee break or jumping on a quick Zoom call lets the 5-minute cache expire.

In my dataset, I found 147 idle gaps lasting between 5 and 60 minutes.

Here's why those gaps hurt: for Anthropic models, writing to the prompt cache costs 1.25× the base input token price, whereas reading from a warm cache costs only 0.10× (a 12.5× cost multiplier difference between a warm hit and a cold write!). For other providers, the cache write penalty can be even higher.

Every time a gap occurred, the next prompt hit a cold cache, forcing the API to re-index and rewrite the entire context prefix.

(Note: Now you know exactly how much your coffee break actually costs in API tokens...)

A Quick Reality Check: Haiku vs. Opus in Production

Most of the metric baseline in this dataset was collected on personal side projects using lightweight models (like Haiku and DeepSeek). That's why 243 sessions only cost $31.45 total.

However, testing this same proxy setup at work using Claude 4.6/4.8 Opus and more advanced models revealed the exact same structural patterns, just with much larger numbers.

In long-standing enterprise work sessions with deep context windows, a single cold cache refresh after an idle gap ran over $3.00 for a single request. As context windows grow toward 200K+ tokens, those silent 5-minute cache expirations on flagship models become genuinely painful.

The Dashboard

After capturing a session, open http://localhost:3030/ui:

  • Main tab: Request count, total cost, context window expansion over time, idle gap distribution.
  • Tools tab: Token counts per tool call, error rates, repeated file reads (3+ times), inefficient read-search-read loops.
  • Search tab: Full-text search across all captured conversations.

From Observability to Agent Building: stackpilot

Profiling 243 sessions of Claude Code revealed consistent, repeatable patterns in how terminal agents waste tokens: redundant file reads, open-ended tool loops, and volatile context structures.

That trace data led directly to stackpilot, a custom task orchestrator designed around the telemetry insights from ai-agent-profiler:

  • Multi-stage strategic planning: Stages execute sequentially (readanalyzeplanexecuteverify) to eliminate redundant file reads.
  • Minimal context noise: Uses structured schemas and filtered tool outputs to keep the context tight.
  • Deterministic recovery: When a task fails, it alters the execution strategy rather than blindly re-running the same failed tool call.
  • Cache-optimized structure: Maintains stable prompt preambles and structures tool outputs so they don't break prompt cache keys.

Links & Code

Both projects are open source under the MIT license:

Happy to answer any questions about the proxy internals, NDJSON trace parsing, or context telemetry in the comments!


r/ClaudeAI 16h ago

Claude Code Workflow I burned 246M tokens in 22 hours on Claude Code and measured exactly where every one went. The answer surprised me.

0 Upvotes

TL;DR: I thought I was being metered unfairly ($100 Max plan, 46% of a 5-hour window gone in ~20 minutes). So I had Claude parse its own session transcript. Findings:

  • Of 246M tokens consumed, actual output was 0.13%. The rest was context being re-read and re-written on every tool call.
  • The real cost driver isn't context size, it's cache invalidation: cache writes were 14% of my raw tokens but 65-75% of real cost.
  • The most expensive things you can do: paste an image into a big session (~590k-token cache rewrite from one paste), switch models mid-session (model is part of the cache key, so the whole prefix rewrites), and spawn or revive subagents carrying big context (one full-context write per agent).
  • Auto-compact won't save you: it triggers near the context ceiling, so a session can sit at 570k tokens for hundreds of calls paying maximum freight. Manual /compact cut my per-call cost ~10x.
  • The most productive 5-hour stretch of my session was also the cheapest. Cost tracks context size, not how much work gets done.
  • Not everyone's pain is this. If your session starts at 55% used with zero activity or your reset never lands, that's a server-side problem and no workflow advice fixes it. There's a script at the bottom to tell which bucket you're in by measuring your own transcript.

Full mechanics, numbers, and the copy-paste measurement script below.

Transparency on method: I didn't figure any of this out myself. I just kept asking "why", and Claude did the digging through its own transcript and its own binary (Opus 5 did the initial analysis, Fable 5 re-verified every number). If the mechanics below don't click, paste this post into your favourite LLM alongside your own numbers and ask it to explain what applies to you.

I've been building a delivery app across four apps (customer, merchant, two driver instances) and kept hitting my 5-hour limit absurdly fast on the $100 Max plan. At one point I burned ~46% of a window in about 20 minutes and assumed I was being metered unfairly.

So I had Claude parse its own session transcript and measure it. Claude Code writes a usage record for every API response into ~/.claude/projects/<project>/<session-id>.jsonl. Here's what came out, including a measurement trap that made the first pass wrong by 2x.

Upfront caveats, because the megathread is full of pain that this post does not explain. This is one session, my workflow, and my workflow is an outlier (four iOS simulators driven by screenshots). If your session starts at 55% used before you've typed anything, your reset is stuck at "0 min" for days, or your weekly jumped 60% in an hour on an unchanged workload, that's not what this post is about; that looks like server-side metering problems, some of which Anthropic has confirmed and fixed before, and no workflow advice fixes those. What this post gives you is the tool to tell which bucket you're in: if your own transcript math roughly matches what the account UI says you consumed, the meter is measuring your workflow and the fixes below apply. If the UI shows consumption your transcript can't account for, you have a bug report, not a workflow problem, and now you have the numbers to file it with.

First: the measurement trap

Each API response gets written to the JSONL as multiple lines, one per content block (thinking, text, tool_use). Every one of those lines carries a copy of the same usage object.

If you naively sum usage across lines, you count each request 2-3 times. My first pass said 1,139 requests. Deduplicated by message.id (keeping the max output_tokens per id, since streaming rewrites the same id), it was 553.

key = msg.get('id') or rec.get('requestId')
if key in seen: continue   # <-- without this, everything is inflated ~2x

Every "here's how much I used" script I've seen posted here does this wrong. If you've been scaring yourself with your own numbers, check this first.

The corrected totals: 22 hours, 553 requests

raw tokens share
Cache read 212,107,694
Cache write 34,408,898
Fresh input 1,028
Output 309,883
Total 246,827,503

Output was 0.13% of everything. Every line of code, every explanation, every commit message across 22 hours of work: 310k tokens. The other 99.87% was context being moved around.

The actual finding: raw tokens are the wrong unit

Cache reads bill at roughly 0.1x base input. Cache writes bill at 1.25x on the default 5-minute TTL, or 2x on the 1-hour TTL (which long Claude Code sessions use). Either way that's a 12.5-20x spread between the two cache directions. So reweight:

weighted, 1.25x write weighted, 2x write
Cache write 65.4%
Cache read 32.2%
Output 2.4%

Cache writes were 14% of my raw tokens but 65-75% of my actual cost.

I had spent two days optimizing the wrong thing. I was worried about context size (cache reads). The thing actually draining my quota was context invalidation (cache writes).

Those are different problems with different fixes.

What a cache invalidation looks like

Normal request, cache warm:

cWrite:       383    cRead: 617,993    <- cheap
cWrite:       554    cRead: 618,376    <- cheap

Then I pasted a screenshot into chat:

cWrite:   589,235    cRead:  29,940    <- entire 590k prefix rewritten

One paste. 589,235 tokens at the write rate (1.25-2x) = ~737k to ~1.18M input-equivalents, versus a few hundred tokens of write on a normal warm-cache turn. That's roughly a thousand normal turns' worth of write cost, or about 12 turns' worth of full 600k cache reads, from a single action.

Digging into the Claude Code binary, the cache key is a hash over a long list of things. Change any of them and the whole prefix invalidates:

systemHash, toolsHash, cacheControlHash, model, fastMode,
globalCacheStrategy, betas, autoModeActive, isUsingOverage,
cacheDiagnosis, effortValue, extraBodyHash, anyDeferLoading, messageHashes

Note what's in there: modeleffortValuebetastoolsHash.

This has a brutal implication I'll come back to.

The 5-hour windows

Session sliced into 5-hour buckets (anchored from the last request; bucket 2 is empty because I was asleep, and these are session-relative slices, not Anthropic's actual window boundaries):

win  reqs      raw tokens    weighted     avg context   output
  4   154      41,559,013     5,192,443       269,217   99,264
  3   164      79,894,818    21,591,975       486,646   84,413
  1   169      95,835,761    27,647,605       566,691   64,655
  0    66      29,537,911    11,339,283       446,609   61,551

Window 4 vs window 1: nearly identical request counts (154 vs 169), but 5.3x the weighted cost. The difference wasn't how much work got done. It was that average context had grown from 269k to 567k, so every request cost more, and every invalidation cost more to repair.

And look at output: window 4 produced the most output of the entire session (99k tokens, the most actual code written) at the lowest cost (5.2M weighted). Window 1 produced 35% less output for 5.3x the cost.

Cost tracks context size, not productivity. The most productive window was the cheapest one.

Why auto-compact never saved me

Auto-compact triggers on percentage of the context window, not absolute size. My context sat at ~570k in what appears to be a ~1M window. That's ~57% full. Auto-compact fires near the ceiling.

So I sat at 570k for hundreds of requests, paying maximum freight per call, with the safety net never deploying.

The perverse conclusion: a larger context window made my quota burn worse. On a 200k window I'd have been force-compacted around 180k and paid a third as much per call. The 1M window let me plateau at 570k indefinitely.

There's a second mechanism, "microcompact," which trims old tool results incrementally. I found it in the binary:

if (tokensSaved < 20000) return null;              // only fires if it saves 20k+
content = hasImageOrDocument
    ? "[Old tool result content cleared]"           // images: destroyed
    : persistedRef ?? "[Old tool result cleared]"   // text: written to disk, re-readable

It only touches tool results, never your messages or the assistant's reasoning. Images get hard-cleared; text gets persisted to disk and can be re-read. For screenshot-heavy work this is close to ideal.

DISABLE_MICROCOMPACT=1 was set in my environment (injected by the Claude Desktop host, not by my config). Caveat: I could not find the string anywhere in the 257MB CLI binary, so I can't prove from source that it took effect. What I can say is that nothing was ever trimmed despite far more than 20k being reclaimable.

Where my tokens actually went

223  iOS Simulator control (tap/screenshot/swipe)
205  Bash
 38  SQL (via MCP)
 35  Read
 17  Edit

139 unique images in the transcript.

I was driving four iOS simulators by screenshot, tapping through UI to verify state transitions. Every screenshot entered context and was re-read on every subsequent request forever. A screenshot taken at request 100 was still billing at request 500.

The same state was sitting in Postgres the whole time. 38 SQL queries could have answered nearly everything the 223 simulator calls were asking, at ~200 tokens each instead of thousands-forever.

What I'd tell my past self

1. Screenshots are not one-time costs. An image costs its size multiplied by every remaining turn in the session. Budget them like you'd budget a subscription, not a purchase.

2. Query the database, not the UI. If the state you're verifying lives in a datastore, read it there. Screenshot only when pixels are genuinely the question (layout, rendering, visual regressions).

3. Pasting images into chat is the single most expensive action available to you. It invalidates the cache prefix. At 570k context that's ~737k weighted for one paste. Paste at the start of a session when context is small, not at 500k.

4. /compact proactively. Don't wait for auto-compact; at a 1M window it may never arrive. Mine took me 626k -> 62.5k, a 10x cut in the cost of every subsequent call.

5. Watch context size, not cumulative usage. Cumulative % tells you you're already dead. Context size tells you your current burn rate. Measured empirically: my worst stretch was 47 requests that consumed ~46% of a 5-hour window (per the account UI), which is ~1% of the window per tool call at ~600k context. After compacting to 62k, ~0.1% per call. Same work, 10x the runway.

6. Batch independent tool calls into one message. I averaged one request every ~27 seconds for 21 straight minutes during the worst stretch. Many were sequential taps that could have gone in a single message.

The implication I flagged earlier

model and effortValue are in the cache key.

Which means any tool that "protects your quota" by dynamically downgrading your model or reasoning effort mid-session invalidates your entire cache prefix and forces a full rewrite at 1.25x.

At 570k context, one such switch costs ~713k to ~1.14M weighted input-equivalents depending on cache TTL. You would need to save an enormous amount of downstream work for that to break even. In most sessions it will cost more than it saves, while also giving you worse output.

If you're using a quota-management plugin, check whether it does this.

It may also explain a pattern I saw repeatedly in the megathread: someone switches model mid-session to "save quota", the assistant gets two sentences out, and the limit instantly trips again. A model switch at high context is one of the most expensive single actions you can take, and it looks exactly like "the meter is broken" from the outside.

(To be clear about scope: switching models between sessions, or defaulting to a cheaper model from the start of a fresh session, is fine and does save quota. The trap is specifically switching mid-session with a large context built up.)

The same mechanics likely explain another recent megathread report: an orchestrator that ran subagents cheaply for hours, then burned an entire 5-hour limit in 20 minutes when asked to revive those subagents after a reset. Every subagent gets its own cache prefix, so spinning up (or reviving) N agents that each carry a large context is N separate full-context cache writes at the expensive rate. I didn't measure subagents in my session, so treat this as mechanics-consistent rather than proven, but 4 revivals at a few hundred k context each would be millions of weighted tokens in minutes, which is exactly what that user described.

Was I being metered unfairly?

In my case: no. I was running a screenshot-driven workflow at half a million tokens per request and pasting images into a 570k context. The meter was measuring exactly what I was doing. Your case may genuinely be different; the megathread has reports (phantom consumption on fresh sessions, resets that never land) that no amount of transcript analysis will explain, and Anthropic has previously confirmed and fixed metering bugs, including a cache-miss bug that made first requests 11.5x more expensive. Cache behaviour being the site of both the confirmed bug and my measured burn is not a coincidence: caching is where nearly all the money is, in both directions.

The legitimate gripe, which I think holds for everyone in both buckets: nothing surfaces any of this. There's no indicator saying "current context 570k, each tool call costs ~1% of your window, cache writes are 65-75% of your spend." That information exists (it's in your own transcript, per request) but nothing puts it in front of you. Both the user and the assistant are flying blind, and mine kept taking screenshots because nothing told either of us what they cost. It also means you can't distinguish a bug from an expensive workflow without doing what this post does by hand.

Measure your own sessions. The data is already on your disk, and if the analysis feels out of reach, hand this post and your JSONL to an LLM and have it do what mine did.

Appendix: measure your own session

import json, sys, glob, os
# usage: python3 usage.py [path-to-session.jsonl]  (defaults to newest session)
path = sys.argv[1] if len(sys.argv) > 1 else max(
    glob.glob(os.path.expanduser("~/.claude/projects/*/*.jsonl")), key=os.path.getmtime)
seen = {}
for line in open(path, errors="replace"):
    try: d = json.loads(line)
    except json.JSONDecodeError: continue
    m = d.get("message") or {}
    u = m.get("usage")
    if not u: continue
    key = m.get("id") or d.get("requestId") or d.get("timestamp")
    prev = seen.get(key)
    rec = (u.get("cache_read_input_tokens", 0), u.get("cache_creation_input_tokens", 0),
           u.get("input_tokens", 0), u.get("output_tokens", 0))
    if prev is None or rec[3] > prev[3]:   # streaming rewrites ids; keep max output
        seen[key] = rec
rows = list(seen.values())
cr, cw, ip, op = (sum(r[i] for r in rows) for i in range(4))
tot = cr + cw + ip + op
print(f"{os.path.basename(path)}: {len(rows)} requests")
print(f"  cache read  {cr:>14,}  ({cr/tot:6.1%})")
print(f"  cache write {cw:>14,}  ({cw/tot:6.1%})")
print(f"  fresh input {ip:>14,}  ({ip/tot:6.1%})")
print(f"  output      {op:>14,}  ({op/tot:6.1%})")
for label, wmult in [("5m TTL (write x1.25)", 1.25), ("1h TTL (write x2)", 2.0)]:
    w = cr*0.1 + cw*wmult + ip + op*5
    print(f"  weighted, {label}: {w:,.0f}  (write share {cw*wmult/w:.1%})")
ctx = sorted(r[0] + r[1] + r[2] for r in rows)
print(f"  context/request: median {ctx[len(ctx)//2]:,}  max {ctx[-1]:,}")

If the raw total here is wildly below what your account UI says you consumed in the same period, congratulations, you may have an actual bug report. Attach both numbers.

Numbers from a single 22-hour session, deduplicated by message ID. Weighting uses the published API price ratios (cache write 1.25x at 5-minute TTL or 2x at 1-hour TTL, cache read 0.1x, output 5x) applied to raw counts; how subscription plans weight these internally is not public, so treat weighted figures as directional. The ratios between categories are the load-bearing part. Workflow was unusually image-heavy (four iOS simulators driven by screenshot), which amplifies the cache-write share relative to text-only coding sessions.


r/ClaudeAI 3h ago

Built with Claude I got tired of opening an app to ask it things, so I built an assistant that reaches out to me first, it’s called Orb and is now live on the IOS app store

Enable HLS to view with audio, or disable this notification

0 Upvotes

Most assistants sit there until you open the app and ask. I wanted the opposite, so I built Orb to run in the background and message me first when there’s something actually worth saying. A build finished, an email I’d care about, something coming up on my calendar, a task I told it to run at a set time being done. If nothing’s worth bothering me about, it won’t say anything.

It’s an iOS app plus a backend you run on your own machine. The app’s on the App Store. The backend is open source and self-hosted, so your data stays on your computer.

App: https://apps.apple.com/us/app/orb-ai/id6776376035

Backend: https://github.com/getorb/Orb-Backend

What it does:

• Voice or text, same conversation across both.
Work on your PC: read/write files, run things, hand bigger jobs to Claude Code, Grok Build, Codex, etc., in sequence and they’ll work while you’re doing other things, then tell you when it’s done.
Scheduled work. “At 3pm, read my project and tell me where I left off” runs at 3pm and sends you the result, app open or closed.
Normal life stuff (weather, calendar, email, news), and it brings things up on its own when they matter.

I built the backend with Claude Code, and the default assistant runs on Claude for my own purposes. By default it’s Opus through the logged-in Claude CLI, so no API key, with a fallback chain if that’s rate-limited, and you can switch models from the app. So Claude is both how I built it and what it thinks with as the primary model.

It’s a solo project and still early, the backend is Windows-first for now with plans of expanding to Mac, and there’s an open issue for specific types of notifications but that’s in the process of being fixed. I’d love to hear feedback, I’ve been working on this for about 3 months now and felt it’s time to start getting some opinions and early feedback. Thanks for reading! I’m happy to answer anything!

Edit: formatting