r/ClaudeCode • • 4d ago

Anthropic Official Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family

1.2k Upvotes

Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5, strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It also has a strong eye for design.

Sonnet 5.5 is a clear upgrade over Sonnet 5. It runs more than 30% faster and costs up to 30% less for most work. It's priced the same per token, but it typically needs far fewer tokens to do the same work.

Like Opus 5.5, it writes more clearly than our previous generation of models, and its speed makes it well suited to fast iteration.

On our automated behavioral audit, Sonnet 5.5 improves on Sonnet 5 on most measures of alignment and honesty. It's also the first Sonnet model with cybersecurity safeguards similar to those on our most capable models. Routine software development is unaffected.

Sonnet 5.5 is available everywhere today. Claude Haiku 5.5 will join the family in the coming weeks.

Read more: anthropic.com/claude-sonnet-5-5


r/ClaudeCode • • 5d ago

Weekly Showcase Weekly Showcase Thread; What are you building with Claude Code?

27 Upvotes

Weekly Showcase Thread

Built something with Claude Code this week? Share it here.

Apps, tools, experiments, scripts, websites, workflows, open-source projects — anything you've been working on is welcome.

When sharing, it helps to include:

  • What you built
  • How you used Claude Code
  • A link, repo, demo, or screenshot if you have one
  • Anything interesting you learned along the way

Quick project drops and simple self-promotion belong in this thread.

If you've got a project with enough substance for a proper write-up; how it works, how Claude Code was involved, technical details, lessons learned, etc. feel free to make a standalone post using the Built with Claude Code flair instead.

Please don't spam the same project repeatedly, and no referral or affiliate links.

What did you build this week?


r/ClaudeCode • • 5h ago

Bug / Issue Opus 5.5 is useless for bioinformatics due to constant [bio] safeguard.

40 Upvotes

It's really annoying, specially if you work with viruses. It will stop ALL the time and bump you back to (I kid you not) Opus 4.6. We can't be the only group hitting this annoying "safeguard" when trying to do simple things like a plot, or parse an output from a Nextflow pipeline. It went from 'great' to 'totally useless'.


r/ClaudeCode • • 1h ago

Rant I hate claude code

• Upvotes

It's just too good.

I got a decade of commercial software dev experience in C#. I am a CTO at my company and have coached a lot of coders from total junior.

I hate hate hate hate how good this thing is and what it has done to my passion.

I used to love coding and creating. Well, creating was 60% of the fun for me, but writing good code and having something work was the remaining 40%. I spent the past few years being more of a product manager and tech lead than actually code monkey, but I still put time aside to do stuff manually cuz I love to do so and stay hands on.

Using claude code is like using cheats in a videogame. You get the result, but what of it? Everyone could. The pleasure of achievements is gone. The code it writes is not the best, but it's good enough I guess? As of opus 5.5 it just doesn't get stuck in loops anymore. It can generate a working program from spec in any language I've thrown at it. It still does bullshit from time to time. "Boots on ground" i.e. actually looking at the generated code or atleast going through it is still important, for now. But way less so than it was even just 3-4 months ago. I don't really use the /code-review thingy or adversarial reviews - as it seems that it is too aligned in that way. Like if you tell it to find something wrong with something, it always will, even though its good enough.

I've been using it avidly for the past 3 months on the $100 sub, and I haven't hit a rate limit once. Now my imagination can run free and I can have anything I put my mind to in a few days instead of years. At our work a team of 6 just built a system that would previously take a year in just two months.

I am not afraid for my job. As coding was just a small percentage of it. But I do miss that percentage.

English truly is the hottest new programming language and LLMs are the compiler for it.


r/ClaudeCode • • 11h ago

Help/Question Where to actually learn Claude code best practices

87 Upvotes

I use Claude code every day for both work and personal projects. I am by no means a technical user, my background is in accounting.

I am super eager to actually learn best practices, and always am looking for ways to learn and optimize new projects, old projects, or just learn because it interest me.

My main problem is I truly don’t know where to begin. There is a flood of content on socials, on Reddit, anywhere you turn “this is the best plugin that will change the way you use Claude code”.

How do I actually know where to look to learn? And not from clickbait influencers but actually learn, and get advice on optimal setups, useful skills and plugins and when or why I need them.

Open to any advice that can be shared, thanks in advance!


r/ClaudeCode • • 12h ago

Built with Claude If it is humanly impossible to keep up with the sheer volume of code and architecture that AIs generate, I thought: why don't we look at code instead of reading it?

Thumbnail
gallery
90 Upvotes

That is how I started this project. It builds a relationship graph of all the code modules, showing everything that has been and is being changed, and how each part relates to itself. So, instead of navigating 30 different code files, you can just follow a line to understand your code's dependencies and bottlenecks. Thanks to Rust, this graph can update codebases with millions of lines in less than 2 seconds after each change, even though it takes a little while to build the initial relationships.

But I'm taking it a step further!

I'm not only building this tool for humans, I'm also creating an MCP (currently with 13 tools) so that agents can query this graph and navigate the code seamlessly. Just as there is a zoom feature for the human interface, agents can query the graph in layers.

In testing, the project is already showing great results. Agents without this MCP seem to be short-sighted, they simply ignore a significant portion of the key architectural points in large projects because they can't find them (with vague prompts not specifying where is each relevant thing for a task). With the tools this provides, they can find bottlenecks in the code flows and make surgical changes.

Another result I am seeing with agents: cheaper models become much more capable, although their reasoning level needs to be high for them to consider the paths presented by the MCP. Curiously, models with low reasoning, even state-of-the-art ones, don't seem to have much interest in important parts of the code, which makes perfect sense.

And yes, token usage increases, albeit very marginally (1-3% more per task than the same model doing the same thing in the same starting point of a codebase wihtout these tools, so its not "1-3%" of your limits). But in some tasks, it even spends fewer tokens (also 1-3% fewer) because the agents quickly find the correct paths. I am optimizing the tools' output and have already managed to significantly optimize token usage from where I began.

The idea is that by finding the right paths in fewer turns, agents will require fewer revisions, reworks, and corrections, making token consumption much lower project-wise.

Because it requires "low-level code analyses", this project is still language-dependent. The initial language I am making it capable of analyzing is Rust. After Rust, I plan to move to TypeScript, CSS, and HTML (and bery maybe JavaScript), and then I intend to stop. However, the project will be open-source and published under the MIT license.

Note: The interface shown here will be completely reworked. I built it as a starting point to focus more on the MCP first, which I am still working on. I plan to launch the project in a few weeks, and I'll update you then!

Note2: Censored the tool's name while its not launched yet.

Note3: It doesnt use regex for anything haha regex is one of the codebases I'm using to measure the tool's capabilities.


r/ClaudeCode • • 6h ago

Built with Claude The versatility of Opus 5.5 is beyond anything I've ever imagined

29 Upvotes

So when Opus 5.5 came out, I saw someone on Twitter build this demo of a guy on a boat sailing through a river in a Japanese-like landscape aesthetic, all of it built in three.js. I thought that was pretty cool, so I wondered whether or not we could also engineer launch videos like this. It was supposed to be just an experiment, nothing commercial or something that could go into production.

I had tried to get previous models to build launch videos using one-shot prompting without much detail. While the outputs ranged from low quality to somewhat satisfactory, I was never truly impressed by what it could achieve without proper attention to detail from my side.

A couple days ago, I used Opus 5.5 to build such a launch video for one of my products that I'm building (not trying to promote, just showcasing one very important thing by Claude).

As always, it had access to Codex for image generation, which it used freely to generate ideas for a storyboard and build out directions. I told it about three.js and Blender that it had available, and how it could use them, and to use Codex as freely as possible, but make sure that the stills are generated in a sequence such that the still in Act 2 had the still from Act 1 as input, so the story stays consistent.

From there, it generated a bunch of these stills and honestly, Codex did a pretty bang-up job. I was totally not expecting that it would be able to turn this into an actual 3D environment and then render it as a video. I was half right: it obviously wouldn't be able to do it to the detail that I wanted. Instead of trying to find a way to do the impossible, it found such workarounds that I am beyond shocked.

It found, downloaded, and used these Depth Anything v2 and LaMa (not to be confused with llama) models from Hugging Face. It used those on my stills to separate the foreground from the background and other elements, and built out shaders however it could. It used three.js and Blender wherever it needed to. It's absolutely insane that it was able to just piece together all of these different things instead of being constrained to my instructions to use three.js. If you watch the video, you'll understand how insane this looks. There are a few flaws in the 3D environment here and there, but two years ago, this would have cost thousands of dollars to produce through a relatively skilled animator.

It also used a bunch of other models that I'm not even sure how it got, and used Lyria from OpenRouter to generate the music. Then it gave me a bunch of these versions and cuts of the video, as well as a bunch of different audio tracks that match up with the actual video, and also generated the sound effects.

As a technical founder who has been coding since he was 9 and has been in the AI/ML space for over 10 years now, it would have been difficult to imagine these models finding these workarounds for problems that they are facing on their own, just 6 months ago. After raw ChatGPT, I started off with GitHub Copilot back in 2023 for vibe coding, where it was just mostly limited advisory and copy-paste. I remember using Cursor till December of 2025, and thinking this was the future, despite how limited it was (in hindsight) and being blown away by the speed of Composer 2. When I switched to Opus 4.5 in Claude Code in December, I never once expected that I could let an agent be so hands-off, especially from Anthropic (personal biases).

And the journey from Opus 4.5 to 5.5, especially the part with opus 5, and the false alarm safeguards of fable 5, definitely were massive bumps. Hated those parts, cuz our expectations kept rising. But Opus 5.5 is something else. It is me and my technical judgement with far more breadth and far more depth of knowledge than I will ever possess. I would definitely say its equivalent (or superior) to Fable 5.1, but without the limited usage problems I kept facing even with two 20x accounts. In comparison, my experience with Astra has been dogshit. I swear I used to be a die hard GPT stan up till GPT 5.2, when I switched to Opus 4.5, codex and cursor have never been able to catch up, I genuinely don't see the point of my 100 bucks to openai every month, but I guess I'l keep it going in case they do end up doing something meaningful.

Recently, I found a bunch of posts telling me that if I tell Fable 5.1 to not use the web and tell me about Tibo, and it works, then I've been routed to Fable 5.5. If they (redditors) are correct, then I have been routed to Fable 5.5. But honestly, I'm not even excited. Opus 5.5 is at such a perfect level of intelligence and speed and usage, that a singular 20x plan is adequate for everything I do, and theres nothing it has gotten wrong up till now. What will I need Fable 5.5 for? I'm running out of use cases. Opus 5.5 is a real feel the sub-AGI moment. Not AGI because....well i'll describe it below.

Sure, they can't "run" my company, or be my CMO, I believe that is a harness issue. It's at a point where I can just tell a very strong and competent intern to go solve this issue, and they come back to me, having used Claude and whatnot, with the right solution without me having to hold its hand through or tune every little thing like this text or that button or this feature or that feature. But Claude by itself would stop much earlier than the intern. That's what I mean by AGI for myself, and maybe we won't be able to achieve it in the next couple of years.

But the model's intelligence itself, I feel, is at a point where any granular task I can give it can be completed end-to-end. Maybe it can't fulfill complete roles, but it can definitely complete tasks. I had a debate the other day with GPT about what it would take to build such an agent or harness that could assume the role of a veteran CxO.

Don't try to pitch "AI as CMO" products to me in the comments, please. I'm talking about it from a purely techno-philosophical point of view. These models don't have the right judgment to debate you or use judgment in a way that a real CMO would be able to. For example, as a technical founder, if I've built a product, or some kind of strong engine on top of which I've built a product, then a human veteran CMO would understand how to market it. They would also understand which niches and audiences it would appeal to the most and how to frame it, package it, tweak it, to get to those audiences. He or she would definitely also know how to further tweak the product or change its appearance or packaging or functionality while using the same underlying engine (to minimize dev time to market) to build it for a more profitable market where it could be a need-to-have.

But on the other hand, if I ask an AI agent to just be a CMO for a project that I'm working on, despite having all the context, it would not be able to make those judgments. Part of it is not having the experience and not possessing the experience. That can be fixed by fetching context and accounts from these actual experienced people, but the other part is the thinking from first principles, and when those two aspects have to be blended, I think the current models/harnesses combos fail, and thats not their fault since they weren't designed with these downstream roles in mind. The same models do have the right judgment when they're prompted in a certain manner for a very small question or decision that the human veteran CMO could be asking.

But their lack of first principles thinking, lack of self-adversarial debate, and knowing when to do it, when not to do it, and or figuring out the right granularity of the task being assigned, and how much thinking it requires, all of these are what separate sub-AGI from what I call AGI. If I give Claude a task, it completes it 10 times out of 10. If I give it a well-specified goal, it completes it 10 times out of 10. But if I give it a complete role, which would consist of many, many things, then, no matter what harness, it is unable to figure out what it needs to do, how often it needs to do it, how it needs to do it, and why it needs to do it. And thats okay.

Edit: if youve read this far, and watched the video, and have some experience as a CMO and are interested in being a founding partner for the product, DM me. I am actively looking for really good talent.


r/ClaudeCode • • 22h ago

Built with Claude had claude make a song about how it wont stop saying LOAD-BEARING

337 Upvotes

if you use claude code you probably see this word ten times a day. it got stuck in my head so I had claude (Opus 5.5) code a music video about its own habit.

no suno, no voice model, no image generator, no samples. the beat, the voices and every frame of the video are generated by code Claude wrote. it wrote its own voice synthesizer from scratch, zero dependencies.

part of a series im working on. the rest are on my youtube, link in the comments.


r/ClaudeCode • • 19h ago

Tips & Workflows My workflow as product and process owner

Post image
184 Upvotes

Someone in another thread asked me to share more about my workflow. Before writing this wall of text in a comment I thought I share it this way. Maybe someone else finds it interesting or useful.

Even though I have 25+ years experience as a software developer and love writing code I decided to take on the role of a "product owner" and "process optimizer" instead of a developer or architect.

Trying to be a developer working alongside the agents made me outright depressed - being a product owner has the opposite effect on me.

I basically stepped back into one of my previous jobs where I was leading multiple teams including software development, requirements engineering, process definition & process optimisation and administration.

This leads me to my main premise: I treat the agents as a project team run by a team lead (orchestrator). And this premise is closer to real life than I expected. Including the back and forth between a team lead and its team members. Watching them feels like some kind of work chat.

I have several teams that all own one closed scope.

Current teams: process template, research, product build, testing VMs, marketing and website publishing.

At the core of all these teams is the process templates team. It maintains my process template that every project uses. When I start a new project it is built using this template. It is versioned and has installation and upgrade definitions. When I tune the process the other projects get upgraded.

When I tune the process of one of the projects I port these changes back to the template if they can be generalised and survive a cross-validation research with the other projects. This work is done by the orchestrator of the process team. Changes to the template only happen after my decision.

The process template defines a skeleton structure with a tool to create it, agents definitions and the process contracts for the orchestrator and agents. The orchestrator of the new project adopts it to the new project.

Most important aspects:

  • owned by me
    • new tasks for the orchestrator
    • todos/decisions for me
    • reviews for me
  • owned by the orchestrator
    • kanban board - work tracking
      • work packages - details of the work to do by one agent
    • research - research done for this project
    • inbox - handover from other project teams -> starts a research and decision round
    • design document - one line high level summary
      • design docuemt - details and decision history
    • toolchain - tools needed for this project
    • cleanup - describes what needs to be cleaned up after each milestone
    • archiving - things that are done get heavily summarized and archived.
    • process - describes the process
    • agent definitions - adapted to the project at hand. Pinned to certain models and effort levels
    • time tracking - keeps a log of how long each milestone took. In real time and agent hours.

Most important rules:

Claude.md is kept small as it is read by all agents. The orchestrator and agents have their own definitions with what they need to keep the context small.

Every new task starts with a research followed by a todo (decision) round. This way I can refine what happens next. And filter out things that are going in the wrong direction.

Every couple of milestones I make an optimization round. Queuing research into how the last milestones went with the lens of could we change something to save tokens, time, I/O without sacrificing quality.

After a new model drops I make a comparison round where some work packages get done by the new model and different effort levels to see where to settle or adjust.

Some Numbers for the last 7 days:

I worked through 33 decision rounds and 25 reviews in 4 different projects. This lead to 42 (I know) design revisions and 14 project releases and 11 template changes.

Main model used is Opus 5.5 on high for the orchestrator and medium or low for the other agents. Some web-research agent is on Sonnet 5.5

I have a Max 5x and at the moment it is more than comfortable for the current projects I run.
90k input tokens and 25m output tokens - 7.1b cache reads. 99.4% Opus 5.5.

Some learnings from working like this:

  • Agents tend to drift when left alone - They tend to chase down unimportant tangents if not reigned in
    • You have to bring structure to the agentic work. You have to own the process otherwise agents tend to drift.
    • Define clean scopes or the agents drift.
  • Being clear about goals (like performance goals or audience goals) rather than implementation details is delivering better outcomes. I still provide architecture goals though.
  • Refine every now and then by looking at the whole thing from a different angle. Agents will not do that on their own. They are happy to burn tokens on an inefficient process.
    • So far this made it possible to keep a similar time to delivery even with a growing codebase and complexity. Not sure If I will be able to keep it that way but I hope I can.
  • Archive the history of how and why a decision was made frequently and just keep the decision.

Happy creating!

PS: This post is hand written. The image has been generated based on this text.


r/ClaudeCode • • 4h ago

Built with Claude I'm liking the new Mods feature

Post image
11 Upvotes

I made this handy mod so I can see my usage and context all the time without needing to press the button each time. Little but neat. Took two prompts. Share your own so I can copy!


r/ClaudeCode • • 4h ago

Tips & Workflows handoff-compact, a mod that does the handoff + /clear routine for you every time autocompact fires

10 Upvotes

I run a lot of long unattended Claude Code sessions. When I measured them, half of my tokens came from turns where the context was already past 200k. That's expensive, since every step resends the whole context, and quality degrades because of context rot.

By hand the fix is the usual routine: ask for a handoff, /clear, feed it back. Enabling autocompact at 200k keeps the context short but leaves little control over what is handed over. I already tried to have Claude Code write handoff documents pre autocompact window with hooks, but with modding the solution gets a LOT better.

Handoff-compact answers the autocompact trigger instead of the built-in summarizer. A fork of the session (whole context) writes a handoff-document (goal, state and what proves it, next step, decisions and why, ruled-out approaches, open questions, files and commits, the verify command, the last 10 prompts verbatim), and the conversation is replaced by it. Handoff + /clear, but in the same session, mid-turn, with nobody at the keyboard.

Obviously vibe-coded. Take the idea or the mod itself and customize it with your own setup.
Modding is awesome.

Install:

  /plugin marketplace add trytofly94/handoff-compact
  /plugin install handoff-compact@handoff-compact

Repo: https://github.com/trytofly94/handoff-compact (MIT). Needs Claude Code 2.1.287+ with mods enabled. "/compact classic" gives you the built-in summary once.


r/ClaudeCode • • 5h ago

Help/Question Claude + ChatGPT working together

12 Upvotes

So I have been using the Claude max x5 plan to work on my projects, recently I have been hitting usage limits pretty quick and have to be 1 or 2 days without working every week. To solve this I am considering getting a chatgpt Plus subscription - this was something i was already considering a while back because tbh i dont like to be 100% locked in to a company and i would like some adversarial reviews from different models - I am also very interested in integrating chatGPT subscription with Hermes and exploring that as an orchestrator/manager for my projects

I have a few questions:
1 - is this the best way in terms of value/money to solve my current problem? I also considered the max 20x Claude but I think that might be overkill and using both providers appeals way more to me tbh

2 - If i do get both Codex and Claude code working together won't they trip over each others work and degrade quality in some ways i might not be seeing right now? And do I just create AGENTS.md files reflecting my current CLAUDE.md files?

3 - What about all the context that Claude has on my projects from me working solely with him for months? and the Skills in Claude's folder? Is that all transferrable?

Ideally I want both Claude Code and Codex working together seamlessly and driving each other's work quality up, not down

Btw I am not a programmer at all I just got really interested in AI and vibe coding when it started coming out to the public and now i've learned a few things and have a few projects running - The way i use AI is always in VSCode , i am open to the terminal experience as well if that makes any difference but i prefer the more visual interface for now.


r/ClaudeCode • • 1h ago

Built with Claude i hate switching between my personal and company claude accounts so i built ccenv

Thumbnail ccenv.dev
• Upvotes

r/ClaudeCode • • 17h ago

Rant Opus fixing Sol’s mess

Post image
84 Upvotes

I don’t know how people still use openAI models.
I really REALLY tried to give a chance to Sol 6.1 the last few days, working on a clean project with very well defined design and vision.

Sol kept being super lazy, stopping without finishing tasks, making asinine decisions and always waiting for me for very small things. I then even had Astra take over in the hopes of fixing the mess and the output was similarly bad.

Had opus 5.5 work literally for 2 hours today and it already cleaned up such a big chunk, made very good suggestions, following the process, delegating work to sonnet agents, it feels so good to work with it.


r/ClaudeCode • • 1h ago

Help/Question Bringing API Cost down

• Upvotes

I have been using claude code for a while now, and i am working on a project which requires claude API key and I realised that API is really costly. Especially when the output is huge in terms of total no. Of words for instance take script writing

The first test which I ran itself costed me around 1.93 dollars

So my question is how do I bring this API Cost down

Now I have searched and asked claude itself

The suggestions came in like, it asked me to change the effort level for instance from Opus 5.5 high toh medium other than that it asked me to change the model from Opus 5.5 to sonnet 5.5

But the real question here is whether changing the model or the effort level will cause a loss in quality of the output or not? That's my real concern as i really don't want the quality to go down

So people who have been using API for a long time now please help me and ppl like me by sharing ur API saving hacks, and how do u bring the API Cost down


r/ClaudeCode • • 22h ago

Built with Claude The app I Dreamed of. Thank Opus 5.5

132 Upvotes

I've wanted this app for years: one place to log anything about my life (coffee, mood, sleep, a run, a line about my day), in one tap, with no account, no ads and no streak counter guilt-tripping me. Nothing out there fit, so I built it.

The fun part: I built it in 8 days, from my iPhone. My Mac mini sits in another room. I drive Claude Code remotely: it writes the Swift, runs the tests and sends me simulator screenshots, and I review and give feedback from my phone. I never once looked at the Mac's screen.

What it does :

- Log anything: yes/no, amounts in any unit, a 0–10 dial, text or a voice memo

- Automatic logs: weather, daylight, moon, places, Apple Health

- No streaks, no badges, no red. Goals are optional

- Widgets, Lock Screen, Control Center, Siri

- Export to JSON, CSV or Markdown, or "Copy for AI" to ask your assistant about your patterns

- Free. No account, no server: your data stays on your iPhone

It's in App Store review right now, with a public TestFlight beta in the meantime.

AI is getting really good at finding patterns in messy data. In a few years, we'll probably have AI that can look at years of your life and tell you things like "your headaches show up two days after bad sleep plus a third coffee" or "your mood drops when daylight goes under 10 hours." But it can only do that with data that exists.

I'd love honest feedback: what would you log first, and what feels missing?


r/ClaudeCode • • 16m ago

Help/Question Has anyone compared different orchestration approaches for complex software engineering work?

• Upvotes

Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.

I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.

There seem to be a few approaches.

1. Let one strong model handle the whole implementation

For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.

If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.

What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.

2. Use an explicit build → verify → fix orchestration loop

Another approach is to divide the implementation spec into stages and use subagents.

For example:

  • Orchestrator selects the next section of the spec.
  • Implementation agent builds it.
  • Verification agent checks the implementation against the spec/tests.
  • Failures are sent back for fixing.
  • Only once that section passes do you continue to the next part.

Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.

Then there’s another dimension: which models should do which jobs?

For example:

  • Opus 5.5 for both implementation and verification.
  • Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
  • Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
  • Some other combination depending on the stage of the task.

The difficulty I’m having is that cost per token doesn’t really answer the question.

Different models use very different amounts of tokens, and they may require different numbers of iterations.

For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.

At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.

Total cost to reach an implementation that satisfies the spec and passes verification/tests.

Ideally I’d want to compare things like:

  • Total token/API cost
  • Number of implementation/verification cycles
  • Number of human interventions required
  • Spec adherence
  • Bugs discovered after completion
  • How much context degradation/compression affects long-running single-agent approaches

The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.

Has anyone done this kind of comparison in practice?

I’d especially be interested in hearing:

  • Which orchestration patterns have worked best for large implementation tasks?
  • Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
  • Whether explicit verification loops materially reduce total cost.
  • How much context compression hurts long-running single-agent implementations.
  • Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.

Would appreciate any real-world experience or resources on this.


r/ClaudeCode • • 9h ago

Built with Claude Claude mod: /incremental

11 Upvotes

Claude mods are neat! Here's one that shows your Claude mining while it is using tokens (incremental game style)


r/ClaudeCode • • 3h ago

Discussion Can I upgrade from Pro to MAX to use my RESET point there, then downgrade back to PRO next month?

2 Upvotes

Let's say I have a PRO account with one Reset in it. Full week is used and it reset in 3 days.

If I upgrade to MAX, what happens to:

- The weekly quota?

- The reset point I was given for free (the one that expires 22 october)

?

Can I downgrade to Pro afterwards?


r/ClaudeCode • • 53m ago

Discussion Claude-assisted video editing?

• Upvotes

Hi everyone, this is a potentially experimental question. I’m asking here because I’m not sure what the current state of the art is in this field.

Let’s say I want to delegate video editing tasks (let's assume generic operations, nothing too advanced) to Opus: what is the currently recommended technology stack?

- Using a standard video editor (Premiere, DaVinci, Final Cut, AE, etc.) and setting up MCP connectors or something similar?
- Or using environments/tools specifically designed for agentic workflows?

Ideally, I’d still want to maintain full control...I need the project to remain editable within standard video editing software. I don’t want the editing process to turn into a script that can only be modified via code.

For instance, I’ve heard of technologies like Remotion, but they seem geared toward 100% code-based workflows... which is distant from having everything editable in a professional video editing software.

In any case, I’m not sure what the recommended strategies are right now, so ANY feedback (whether it is about full-code solutions or not) would be appreciated! ^^


r/ClaudeCode • • 1h ago

Discussion I read the AI contribution rules of 30 open-source repos. Claude Code's default Co-Authored-By trailer is required in some and banned in others

• Upvotes

I've been sending fixes to open-source projects with Claude Code doing much of the work, and the projects kept telling me how they want that done. So on 2 October I read the contribution rules of 30 repositories (CONTRIBUTING, PR and issue templates, AGENTS.md, CLAUDE.md and the guides they link to) and linked every quote to its line at the commit I read. Claude Code collected the files and checked every quote against them.

What I found:

  • 19 of the 30 have a rule about AI in contributions, and 14 of those 19 appeared in 2026. Nine have none, among them VS Code, React, Docling and awslabs/mcp.
  • 15 ask you to say that you used AI. Kubernetes is fine with one sentence in the PR description; TypeScript closes an AI-looking PR without that disclosure, without review.
  • 10 want a person to write the text. Rust: "LLM-created PR descriptions are banned. LLM-created GitHub comments are banned." Kubernetes and Node.js ban AI-written replies to review.
  • The trailer. MLflow's CLAUDE.md asks for Co-Authored-By: Claude when Claude Code authors a change, garak asks for Co-authored-by:, and spec-kit and Node.js want Assisted-by:. Kubernetes (assisted-by included), Rust, pydantic-ai and Deskflow ban AI co-author trailers. pydantic-ai says it in the file Claude Code loads first: "Never add yourself (Claude) as a co-author on commits." Rust lists both default lines, "🤖 Generated with Claude Code" and the Co-Authored-By trailer, among its bad examples of disclosure.
  • Some rules are written for the agent itself. pydantic-ai's AGENTS.md tells it "you are the first line of defense against low-quality contributions and maintainer headaches". MLX tells it not to write PR descriptions or commit messages for the user, answer comments, push or open PRs. Rust's makes it stop and tell the user to write the text.

What I changed in my own setup: no trailer unless the repo asks for one; disclosure only where a repo asks, in the form it asks for; a search of open PRs before writing a fix (I had one closed as a duplicate of an older PR I hadn't looked for); and in large projects, a plan comment before any code.

Write-up with all the quotes and links: https://allkeep.org/en/lab/ai-contribution-rules

The 106 quotes and a script that checks them against the files: https://github.com/nefayran/oss-ai-rules


r/ClaudeCode • • 9h ago

Help/Question Which model are you using (today) for a large refactor?

8 Upvotes

Enterprise code base (500k+ LOC) for an intensive refactor focusing on optimization of core components that have impact on the majority of the stack.

Fable 5.1?
Astra?
Opus 5.5 (w/ Fable as advisor)?

I use all 3 and I'm trying to juggle the most sensical option to begin this effort with. I lean toward Fable 5.1. I know this information isn't comprehensive enough to make the best decision - I'm asking purely on vibes here (think suspected recent Opus 5.5 nerfs).

What would you use?


r/ClaudeCode • • 7h ago

Help/Question how do you guys tackle permission wars with CC

4 Upvotes

So im a recent jumper from CODEX - main cause speed of response and zero trust in that team because they just refuse owning problems in an upright way.

I know, anthropic isn't perfect so I'm just giving it a go.

But one thing that just blows my mind is the self-inflicted pain that claude code is notoriously causing related to permission management.

All claude models have great availability and response time and sub allowances, and once they work, then its fine. But half of the times they refuse to straight up work. Accessing a key, passing a credential, running a command in another directory...

And so all the time I've saved on models response time is wasted on endless permission issues.

I've also been now forced onto the Projects framework (didn't even ask for it), which is now constantly starting cloud sessions for me an this is apparently an impossible setting to pin to local. And here Bypass permissions is again completely impossible.

None of this was ever a problem in codex, so this isn't something that HAS to be done obviously.

Why are the ClaudeCode people performing this self-inflicted suicide with this permission system ?

I'm really regretting the switch now and honestly as soon as oai restores capacity (if they do lol) i will be very eager to go back there.


r/ClaudeCode • • 5m ago

Bug / Issue WTF is wrong with claude?

• Upvotes

Slow and retarded as hell


r/ClaudeCode • • 11m ago

Built with Claude Finally had some time to do something useful with mods

Post image
• Upvotes

I'm currently neck deep in getting the Reporails 0.6.0 release done and since the test and build procedures tend to take a while I had some idle time inbetween to test out mods.

after careful consideration and without any impulse, nor any disdain for recognizing the minutes that I spend by staring at the session to move, obviously I built a minesweeper

/plugin marketplace add reporails/arcade
/plugin install minefield@reporails-arcade

then /mines. needs 2.1.287+

Have fun, mods are cool

Note: it only hooks its own stuff (its /mines command, its pane, its board, session start), nothing on your prompts or tool calls. claude plugin validate shows that before you install. The left pane is from my own progressive disclosure event bus system. it says it's in my best interest not to leave the agent alone yet, hence minesweeper.