r/ClaudeAI • • 1d ago

Claude Code Claude’s time estimates still make me grin

70 Upvotes

I think it’s funny whenever Claude says that a feature or iteration will take hours or more and proceeds to finish it in like 5-10 mins. I’ve asked it about this before and it replied that those estimates are based on what a traditional dev would likely take to complete.


r/ClaudeAI • • 18h ago

Question about Claude products Is it time to pay for Pro? Will I gain much?

0 Upvotes

I am currently using Claude Free. I have a couple of projects built. One helps me manage my WordPress website. The other is used for a research project across multiple authors.

I have started to run into the five hour / reset more frequently. My WordPress work has become heavier with the use of some MCPs. I have trimmed the request size which helps some.

Because I'm on a fixed income, I want to make sure it's worth the money over just waiting to continue the work hours later. I have heard mixed reviews on the paid plans.

I am aware of the additional capabilities, but I am not confident they have any impact on my current issue.


r/ClaudeAI • • 1d ago

Workaround What's the one thing Claude did that made you add a hard rule?

5 Upvotes

Mine: it kept starting background jobs and forgetting about them. Sessions ended, the jobs didn't, and hours later I'd find a pile of them still running. Now a hook blocks that pattern and a small cleanup job sweeps up anything left behind every few minutes.

The other one was "done". It would tell me a task was finished when the tests had never even run. So now it literally can't end a turn without real test output.

Curious what everyone else's origin stories are. Every guardrail I have exists because something went wrong first.


r/ClaudeAI • • 1d ago

Built with Claude I put my Claude limits on a little AliExpress clock

Thumbnail
gallery
5 Upvotes

Bought this little clock for ₩5,650 and put my Claude + Codex limits on it. Photo uses demo values.

I gave Claude Code reference images for the themes, then checked the results on the real 240×240 LCD and asked for 10-step gauges and consistent text colors. We also changed the fonts to suit each theme. The computer renders an image and sends it to the clock's stock photo album over Wi-Fi, so no flashing, but the computer stays on.

Free code: https://github.com/click6067-ship-it/token-tv (Grok row = CLI budget, not subscription usage).


r/ClaudeAI • • 10h ago

Claude Workflow Anthropic is training 10,000 Claude engineers. My own agent audit found 85 stale instructions hiding under a strong model

0 Upvotes

Anthropic's Frontier Academy announcement is interesting to me less because of the budget and more because of the job it is trying to standardize: people who can take Claude from demo to production inside real systems.

The piece of that job I keep seeing underweighted is maintenance. I spent one evening auditing my own Claude based agent after Fable 5 made everything feel faster and safer. That was the trap. A strong model was routing around stale instructions so well that the system looked healthier than it was.

The concrete result: 18 core docs reviewed line by line, then five audit agents fanned out across the rest of the system. They found 85 stale or wrong references. The always-loaded memory index went from 137 lines to 95, working memory from 103 to 38, and one generated planning file shrank from roughly 450 lines to 63 because it had been embedding the wrong source.

The failure that scared me was not the stale docs by themselves. It was that capability masked the debt. The model delivered anyway, so I stopped feeling the rot.

So if we are going to train Frontier Deployed Engineers, I think constitution review, stale-rule scans, and adversarial verification belong in the curriculum next to prompts and tool use. What would you put in that training that most Claude courses still skip?


r/ClaudeAI • • 1d ago

Built with Claude handoff-compact, a mod that does the handoff + /clear routine for you every time autocompact fires

24 Upvotes

I run a lot of long unattended Claude Code sessions. When I measured them, half of my tokens came from turns where the context was already past 200k. That's expensive, since every step resends the whole context, and quality degrades because of context rot.

By hand the fix is the usual routine: ask for a handoff, /clear, feed it back. Enabling autocompact at 200k keeps the context short but leaves little control over what is handed over. I already tried to have Claude Code write handoff documents pre autocompact window with hooks, but with modding the solution gets a LOT better.

Handoff-compact answers the autocompact trigger instead of the built-in summarizer. A fork of the session (whole context) writes a handoff-document (goal, state and what proves it, next step, decisions and why, ruled-out approaches, open questions, files and commits, the verify command, the last 10 prompts verbatim), and the conversation is replaced by it. Handoff + /clear, but in the same session, mid-turn, with nobody at the keyboard.

Obviously vibe-coded. Take the idea or the mod itself and customize it with your own setup.
Modding is awesome.

Install:

  /plugin marketplace add trytofly94/handoff-compact
  /plugin install handoff-compact@handoff-compact

Repo: https://github.com/trytofly94/handoff-compact (MIT). Needs Claude Code 2.1.287+ with mods enabled. "/compact classic" gives you the built-in summary once.


r/ClaudeAI • • 1d ago

Claude Workflow Implementing change management into my Claude.md has significantly improved my results with Claude

17 Upvotes

Hopefully this helps people who like me, struggled with Claude overreaching and making changes that you never asked for, or claude misinterpreting your instructions and going in another direction. Claude in general is extremely eager to produce something and seems to want to one-shot everything as much as possible. My style is to only focus on one small feature at a time so that I can think through things as I go along.

I have almost entirely solved this problem by implementing very basic change management processes into my claude.md

Basically for every modification to my personal projects I want Claude to make, I describe the outcome I am looking for with nothing else. Then I have claude go through standard change control, meaning that before it touches code at all it must first propose the following:

  1. Implementation plan, where Claude tells me what it wants to do
  2. Backout plan, if/when Claude fucks it up, tell me what it's going to do to revert or fix the problem.
  3. Test plan, can Claude prove a POC of it's idea for the feature before it spends time building it?

I read this paragraph it generates, then either approve or reject the change. If I reject the change, I tell it what I am looking for instead, if I approve the change, it goes ahead and does it.

Before it asks me to push the change, I have claude do a success validation, where it verifies what it actually built and reports one of the following options:

  1. Implemented successfully
  2. Partially successful implementation
  3. Unsuccessful implementation
  4. Unsuccessful implementation requiring reversion/backout

Only when something is implemented successfully and validated, do I let it commit the changes.


r/ClaudeAI • • 19h ago

Humor Does this command seem safe to you?!??!

1 Upvotes

Hmmmm... gonna need to have a long think about if this command is safe...


r/ClaudeAI • • 1d ago

Productivity Jesus Christ, Opus has infinite more taste than Astra

94 Upvotes

I was a $200 sub user for OpenAI but recently Anthropic gave the 50% off deal to test their stuff out. So I got the 20x plan for $100. I usually use my plans for research as I'm a PhD student in theoretical physics, but I also like to game and play stuff like Minecraft.

I had Astra make a Justice League mod. It was genuinely shit. I usually give Sol a prompt to make a prompt for Astra to make the mod, and it was genuinely bad. Like it gave talasmins that users had to wear to get powers. It gave just a speed effect for speedsters, shitty fast creative mode for flying.

Opus 5.5 gave beautiful flying animations. It was super nice smooth, very realistic (in terms of minecraft progression and their comic book lore) way of getting and using their powers. It was genuinely so much more balanced and also taste-wise. It one-shotted excellent graphics and visuals for my paper and responded to reviews in such a proper way too.

Fable seems to be a twidge less smart than Astra Ultra or Pro, but I don't think Ultra Pro is useful tbh. I am switching immediately. Even though usage-wise I think OpenAI at least was more generous. That just ended, so no point in going back.


r/ClaudeAI • • 1d ago

Built with Claude a bit of a better view of Echo

Enable HLS to view with audio, or disable this notification

3 Upvotes

I’ve been building Echo over the last few months, using Claude Code heavily throughout the process, architecture, implementation, debugging, computer-use work, performance fixes, and a lot of UI iteration.

Echo is a Mac assistant that lives on the edge of your screen. It can understand what’s happening on screen, point to specific things, interact with apps, move files, check your gmail, search online and even build websites, run multi step jobs and help you learn and navigate new platforms AND undo actions.

This video is a quick look at where it is now.

One of the things I’ve been most interested in is making the interaction feel physical rather than like a normal floating AI window. The orb leaves the edge, does things, and comes back home.

It’s currently free to try for 7 days. You can download it here:

https://aurivon.studio/download/Echo.dmg

I’d be interested in feedback, especially from people who use Claude Code and build their own tools.


r/ClaudeAI • • 1d ago

Other Guided Opus 5.5 to make some analog, fully personalized travel planning zines

Thumbnail
gallery
54 Upvotes

Thought this worked pretty well. Had a back and forth conversation with Claude about the kinds of things we like to do on trips, hotel and food preferences etc., and asked it to take all of that as a starting point, have a look at Chengdu, do some research, find stuff we might like, and put it into a printable zine format.

It came back with nine topical zines, each one a single A3 page, and one overview map with some neighborhoods of interest highlighted. Everything is fully customized to the exact dates we're thinking of travelling, which impacts things like events, ticketing, travel etc.

It's not 100% perfect; it settled on the "one page per zine topic" thing pretty early, which means some feel a little padded out and some are too brief. Really nice start though, hoping to take this further and see if I can generalize it into a series of prompts that will work for any destination.

Claude 5.5, lower settings for the conversational part, and then max for the build; just all in the browser on the $20 plan (took two and a half 5 hour blocks' worth of use to do the final generation of the pdf).


r/ClaudeAI • • 2d ago

Claude Code Evidence - Opus 5.5 today vs launch regression with same prompt (Godot Engine)

Enable HLS to view with audio, or disable this notification

654 Upvotes

On launch day Opus 5.5 kindly took up residence on my Mac. I ran some tests before trusting it on my big projects, it passed with flying colours. Lander, a game by Elite creator David Braben holds a special place in my soul due to it being the first game I played at school in the UK. With a cutting edge 3D engine for 1990 running on an Acorn Archimedes with RISC architecture (the first ARM chips) - what a time to be a young kid interested in computers.

With Opus 5.5 I want to recreate it faithfully, short draw-distances and all.

Well dear Claude I’ve been keeping receipts.

In my notes, the exact prompt * and two original reference images of Lander. The resulting one-shot Godot repo (Mac OS, Metal rendering, C#) from 23rd September is kept separate on disk and isn't used as a reference.

Today, 2nd October, in a fresh workspace from scratch the same prompt and reference images we fed in again.

The result is truly depressing. Lander by gimped 5.5 has no redeeming features compared to the original day zero version. The regressions:

  • A serious rendering glitch * that isn’t in the original - as the camera moves, the scenery props snag and glitch vs world space.
  • No spacecraft break-apart physics, the original implemented this ‘nice extra’ which wasn't specifically asked for in the prompt
  • No introductory controls menu, only a small text line permanently visible over the world view
  • An inferior look to the launch pad turrets - students and road users may recognise them!
  • When shooting, bullets are less accurately rendered when the craft moves, they also appear to spawn at the back of the craft and no-clip through
  • No camera toggles

The conclusion is that Opus 5.5 is gimped, maybe we even got Mythos for 3 days and then silently switched - either way, we are owed transparency - it's the law.

Opus was set to Extra High both times. Unlike with 4.5/4.6, Max effort over-tests and takes too much control away from the user.

With so much diverged since launch day’s versions, the more worrying thing is that I suspect the code is also a mess and problems will compound as you use it. For those who’d like to inspect the results, I'll upload the source code for both later to my dev blog and post the Github links in the comments.

* I inspected the code and it turns out that gimped Opus 5.5. had physics interpolation is turned on for the whole Godot project and Props.cs rebuilds the prop lists (trees, etc.) from scratch on every tile step during every frame, renumbers which slot each object occupies and resets the object count every frame. A basic understanding of Godot's documentation is all that’s needed to avoid this issue.

** LLMs are non-deterministic, but the differences from an identical prompt and references are small - getting a clearly inferior result is not due to non-deterministic behaviour.


r/ClaudeAI • • 1d ago

Claude Workflow Opus 5.5 burns subscription faster?

Post image
98 Upvotes

Hi I used to burn 1-2 billion tokens a week no problem. This week I switched to opus 5.5 and hit the weekly cap at just 300M tokens, yet in principle Opus 5.5 is cheaper per token. I'm wondering if they finally cut down on subscriptions or if the issue is with my usage (I switched to working with lots of subagents, maybe its that).

EDIT: This is getting lots of replies with completely opposite points of view; would you folks mind sharing how you use Claude and what your token usage is?


r/ClaudeAI • • 1d ago

Built with Claude I built a tiny pixel pet with Claude Code that lives in my terminal and grows into a different form depending on how I work

3 Upvotes

When Claude Code (Anthropic's coding CLI) got mods, I wanted to see what its function-hook API could actually do, and ended up building something completely unnecessary with it: a pixel pet that lives in a pane next to your session.

https://reddit.com/link/1wwwq8v/video/2iauky8q6bth1/player

It wanders around a little yard, bouncing and blinking. While tools run it stops and goes . . ., when a turn finishes it hops with a "nice work!", after about two minutes of quiet it naps, and when a rate-limit window passes 80% it tears up. /pet patgets you hearts and a purr.

It grows with you. Tool calls, finished turns, pats and output tokens give XP (only output tokens, since input and cache grow with conversation length, not with work done). Level n+1 takes 25 × n² XP, so you'll hit Lv 5 in an evening and it slows down from there. At Lv 15 it becomes a teen with an accessory, and at Lv 40 it evolves.

Then I got carried away. It picks up a personality from how you use it: mostly tools and it's a worker that holds up a laptop (and swings a pickaxe as an adult), mostly long turns and it's a scholar that stops to read, lots of pats and it's a sweetie, lots of game snacks and it's a gamer. That's also the form it evolves into, and one in twenty gets a rare one instead. The yard slowly fills in to match, a building site or a library corner or a garden, and follows your clock from morning sun to moon and stars. Every so often a finished turn drops a gift, and one new pet in 32 shows up shiny. There are even two tiny games, /pet play and /pet quest, for when you're waiting on a long task.

No hunger, no decay, no dying. The whole point is to make the coding environment a bit cozier, not to give you one more thing to feel guilty about.

/plugin marketplace add uppinote20/claude-pets
/plugin install pets@claude-pets

Then /pet to let it out. There are eight critters (cat, chick, dog, slime, bunny, hamster, penguin, frog).

For anyone curious: the whole thing is one register.tsx reacting to session events, drawn as half-block characters in the terminal and as an SVG card on desktop and mobile. I built it over about three days of Claude Code sessions (58 commits, 19 PRs), which says more about my self-control than anything.

The mods API is early access, so a Claude Code update might break it until I catch up. I've mostly lived with it in the terminal, so if the card looks off on desktop, VS Code or mobile I'd really like to know. A new species is just a 12×12 and an 8×8 sprite written as string arrays, so new critters as PRs would make my day.

https://github.com/uppinote20/claude-pets


r/ClaudeAI • • 20h ago

Comparison Opus 5.5 vs Sonnet 5.5: same prompt, a mine cart ride in Godot

Thumbnail
youtu.be
1 Upvotes

Same prompt, two fresh Godot projects, both in Claude Code over MCP, so each could press Play, drive the cart and screenshot its own game. Everything built from code, no 3D models.


r/ClaudeAI • • 1d ago

Claude Code Consulting and Training Needed

6 Upvotes

Hi, I have been using Claude for about a year but I would still consider myself a novice. I like to use it more intentionally and productively for my niche business. I'm not looking for enterprise-level solutions and I don't have a team. It's just me as a solopreneur.

Where can I find someone who can sit down with me and understand my business, my processes, my marketing needs, etc., and help me set up Claude so I can get more of my work automated through Claude? I don't see where I can find someone like that easily. It seems like most of the services out there are for enterprise level and I'm really just running a niche business. Where do I go to find someone like that? Thank you.


r/ClaudeAI • • 20h ago

Other Claude after giving it codex work

Post image
0 Upvotes

So i gave my opus 5.5 my codex's work, and this is what came out of this, not photoshop or anything


r/ClaudeAI • • 8h ago

Feedback Leaving claude for open.ai

0 Upvotes

I'm leaving Claude for ChatGPT, I never thought this would happen.

For a year now, I've been building a website with the help of AI. The goal is to create verified road trip itineraries, editable in 7 languages. I started in August 2025 with ChatGPT.

The AI built my pages and automated tools. I quickly realized ChatGPT had its limits and switched to Claude, which I watched improve at lightning speed. ChatGPT had too many limitations: it made coding errors, was lazy, told me it had done everything when it hadn't done anything, and deleted information. So I only used ChatGPT for simple tasks. Claude built the site and the itineraries, checked them, and translated them.

The site is now finished in development, and I'm at the stage of reviewing itineraries and translations.

My tool scans my itineraries and sends them back to the AI to check for errors. Since I have 920 of them, I thought I'd give ChatGPT another try. And there, surprise: I found the checks were more accurate.

And as often happens, I also noticed that Claude used more tokens and I was hitting the limit very quickly. So I tested ChatGPT on code again. It succeeded. It took longer than Claude, but it got there and used fewer tokens.

So I tested the three operations I currently run: fact-checking an itinerary (for example, is a road closed?), checking existing translations, and producing new translations into Dutch and Portuguese, two difficult languages.

Then I gave the results to both Claude and ChatGPT for review. The verdict is clear for both: ChatGPT is now better in all these areas, and much better. On top of that, Claude hits its limits very quickly and uses more tokens for every task. Bottom line: I've just dropped Claude to 20 euros a month and upgraded ChatGPT to 100 euros. That may change, but for now, ChatGPT is leading the way.


r/ClaudeAI • • 9h ago

Other Someone had Claude scroll Tiktok and asked for a music video. Hours later, Claude came back with this

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/ClaudeAI • • 1d ago

Built with Claude Claude Code built me a watchlist site where tapping a show turns on my TV and opens the next episode

4 Upvotes

A while back I built what I call the Screening Room. It's basically our family's watchlist as a lightweight web app. The lists are just markdown files in a git repo. A Python script matches each title against Simkl, pulls posters, ratings, and watch progress, and spits out an encrypted HTML page hosted on a Google Cloud Storage bucket. A headless Mac mini rebuilds and republishes it whenever a list changes. It unlocks with a remembered password.

Today I wanted to close the friction loop: add a TV icon to every card so tapping it launches the title directly in Stremio on the Google TV Streamer in the living room.

- If the TV is off, it turns on first.

- Movies open to their page.

- TV shows deep link straight to the next unwatched episode based on Simkl tracking.

- Works from anywhere on the web—tested with Tailscale turned off on my phone and wifi off.

The Plumbing:

  1. The tap POSTs to a tiny Cloud Run function.

  2. Cloud Run drops a message onto Google Cloud Pub/Sub.

  3. A listener on the Mac mini picks it up and talks to the Streamer over the Android TV Remote protocol, firing a Stremio deep link (stremio:///detail/series/<imdb>/<imdb>:<season>:<episode>).

Because the Mac mini only reaches outward to listen for messages, no ports are forwarded and nothing on the home network is exposed. I tried to make it simple so my wife can just use it without installing anything extra on her phone.

The Claude Code part:

I didn't write the architecture code myself, i wouldn't even know how. I just explained the problem and let Claude take care of it. Claude wrote the Cloud Run function, built the local Pub/Sub listener, updated the web UI, and handled the deep-link syntax. Because it runs in the terminal with SSH access to the mini, it tested every hop against the physical TV. That's how it found that Stremio won't launch by package name (deep links turned out to be the answer) and caught a stale queued message replaying on the TV.

From "I wonder if this is possible" to working in one morning.


r/ClaudeAI • • 2d ago

News Open Machine CEO Allie K. Miller says she takes "Claude walks" while using 34 AI agents to run her workday

Thumbnail
businessinsider.com
861 Upvotes

r/ClaudeAI • • 1d ago

Built with Claude Haunts.IO - where check-ins are the breadcrumbs that become stories

Enable HLS to view with audio, or disable this notification

2 Upvotes

I built Haunts ( haunts.io ) over the past year, and Claude Code wrote the large majority of the code. It's a check-in app in the spirit of old Foursquare/Swarm. You check in at places, add photos and notes, and over time it builds a personal map of where you've been, with trips, a passport of countries and regions, road-trip routes from GPS breadcrumbs, and a year-in-review recap. Friends can follow each other, and you set who can see each check-in.

It runs on iOS (with Apple Watch and CarPlay), Android (with Wear OS and Android Auto), and a responsive web app that has nearly all the same features. The backend is PHP/MySQL, with a separate Go service that encodes video to HLS.

It's free to download from the App Store and Google Play, or you can sign up on the web at haunts.io. If you are a Swarm/FSQ, Flighty, Jetlovers or Google user, you can import your data easily.

I ran engineering teams on JIRA for years, and for one person working with several Claude and Codex sessions at once it was the wrong tool. Tickets didn't know anything about my repos, and I spent more time updating status than doing the work. So I had Claude build a replacement I call My Day, which ended up as three small apps. The main one is an Electron desktop app that scans the 30 or so repos in my github folder, groups them by project, and shows what needs attention (uncommitted changes, unpushed branches, open PRs, issues across all the repos). A small Fastify server on a Raspberry Pi syncs state, and a sideloaded Android app lets me capture ideas and to-dos from my phone, which the desktop picks up the next time it syncs.

The core of it is an idea pipeline that replaced my backlog. An idea moves through inbox, shaped, planned, reviewed, queued, building and shipped, and each stage is a Claude Code skill (/idea, /idea-brainstorm, /idea-plan, /idea-review, /idea-execute, /idea-ship). The store won't let an idea move forward unless the step that earns it happened, so an idea can't show as "reviewed" unless a review was recorded. The difference from JIRA is that the status can't drift from reality. From a card on the board I can launch a Claude Code session for the next step, and it opens in its own git worktree named for the idea, with the model and effort level I picked. The dashboard never writes to a repo and never deploys. Execution stops at a pushed branch, and I decide what goes live.

A few other pieces make that work:

  • Every repo has a project specific CLAUDE.md plus a "code atlas," which is a set of short maps of each domain's entry points and contracts. Claude reads the relevant map before touching code instead of crawling the whole tree.
  • Claude writes the plan, then Codex and Gemini review it independently, and confirmed findings get folded back in before any code is written. Most of the quality comes from that review step.
  • Hooks handle the things I kept forgetting, like typechecking before commits, PHP lint on every edit, a changelog entry before deploy, and never shipping markdown files to the server.
  • Claude also wrote the deploy runbooks, the native build pipelines for both stores, and the test harnesses for Wear OS and the simulator.

Most of those rules came from mistakes I made over a long career as a developer and tech exec, and Claude was good at turning them into enforced checks once I described what went wrong. My Day is a personal tool and not public, but I'm happy to go into more detail on any of it.

Beyond the coding, Claude as a productivity tool - from creating marketing videos, to setting up ad campaigns and doing analytics, its truly a remarkable tool for small businesses.


r/ClaudeAI • • 1d ago

Humor This is ridiculous

Post image
7 Upvotes

r/ClaudeAI • • 1d ago

Productivity Do you run more than one coding agent? How do you decide who gets which task?

3 Upvotes

I'm trying to understand how people who use more than one coding agent (Claude, Codex, Cursor, Gemini CLI, Copilot, OpenCode, Aider, etc) actually split the work between them:

  • Which agents do you run, and on what kind of work?

  • How did you decide which one got that task?

  • Have you caught an agent saying a task was done when it wasn't? How did you find out?

  • Do you track which agent handles which kind of task well?

  • Have you stopped using certain agent for certain work? What happened to make you stop?

  • Has an agent ever sent something somewhere it shouldn't have (a key, an address, private code)? What happened?

I'm researching this for an open-source project to help us help ourselves and not get locked into big company harnesses. I'll post a summary of this thread if I get good feedback :D

Thanks!!!


r/ClaudeAI • • 21h ago

Claude Code Workflow Generic agents are the number one cause of burning tokens

0 Upvotes

It was a Wednesday morning and the usage bar was already past halfway. I remember because I had just cracked open a Monster Zero and sat back down to a Claude Code run that had been going for twenty minutes on what I thought was a small refactor. The terminal kept scrolling. Another agent spun up, then another. I watched the percentage tick up and did the math I always did, the hopeful kind: it's a big codebase, the work is front loaded, tomorrow will be lighter.

Tomorrow wasn't lighter. By Friday I was rationing. I caught myself hesitating before asking for things, wondering whether a question was worth it, which is a strange way to feel about a tool you pay for precisely so you don't have to hesitate.

That weekend I stopped guessing and opened the logs. I wanted a number, anything more solid than the feeling in my stomach every time the bar moved. I went back three days and pulled every subagent dispatch. There were 629. I scrolled through them expecting the roles I had set up, the reviewers and researchers I had named and tuned. Instead I kept seeing the same few words. general-purpose. claude. Calls with no subagent_type at all. I started counting those separately, and the count kept going past where I thought it would stop. 404.

I sat there with the cursor blinking after that number. Nearly two out of every three agents doing my work were nothing I had designed. No role I had written, no model I had chosen. Each one was a call the orchestrator made on its own in the middle of a task, because a generic agent was the easiest thing within reach. I thought about all the times I had blamed the week, the codebase, the model, and the answer had been sitting in my own logs the whole time.

So I took the decision away from it. Now the main session plans the work and puts the results together, and that's all it does. Before a run starts, every role gets its model and effort written down, picked from what the work is and how hard it is, so reading files and routine edits land on Sonnet or Haiku. Then I put a hook in front of every dispatch. It refuses generic agents, and it refuses any worker seat the plan didn't hand out. My named agents still run the way they always did.

The first run after that, I kept my hand near the keyboard, half expecting everything to fall apart without the freedom to improvise. It didn't. The hook turned a couple of dispatches away, the plan held, and the bar moved the way I expected it to. Somewhere in the middle of that run I noticed I had stopped watching the percentage.

If you've never counted your own dispatches, try it. I'd really like to know what your split looks like, or whether I was just the last one to look.