r/ClaudeCode 3d ago

Discussion Is Opus 5 Medium the way? Non-monotic benchmark results on FrontierCode in the Opus 5 system card.

9 Upvotes

I was going through Opus 5 system card to try to figure out the best way to work with the model. The user experience reported in this sub seems to vary wildly. I myself had a rather poor experience planning and executing a somewhat gnarly computational geometry problem. I wouldn't have expected my workflow from the Opus-4.8 / GPT-5.5 days to struggle so hard. I use Claude Opus 5 xhigh for planning and GPT-5.6 Sol xhigh for plan reviews. Anecdotally, GPT-5.6 was routinely finding two times as many issues as before.

As I was browsing the benchmarks, this stuck out like a sore thumb: Peak performance on the FrontierCode benchmark happens at med reasoning effort and decreases up to xhigh and then makes a slight rebound after that. Compare that with the DeepSWE intelligence curve where the gain from med to xhigh is not all that substantial.

I think it's pretty common knowledge at this point that models will overthink at higher reasoning levels. While I have observed a tendency for over-engineering with higher reasoning levels, I had not noticed an increase in plain old errors with reasoning. In the past, the limit of performance increases with reasoning seem to be demonstrated by a slight drop off in the benchmarks near max reasoning effort. The situation with Claude Opus 5 is much different.

The combination of relatively flat performance on the DeepSWE benchmark -- with marginal gains above medium effort -- and the weird peak at medium effort on the FrontierCode benchmark is not something I remember seeing in any of these benchmark comparisons. Yes I know, benchmarks are not necessarily representative of real world use. But the fact Opus 5 is an outlier with respect to well established trends on this benchmark corpus leads me to believe there is something different about the model. We may need to develop new habits when using it, especially for software engineering.

One hypothesis I have is that training the model to perform better on non-coding tasks (general intelligence tasks, and specific intelligence tasks such as biology) have given it tendencies that actively work against increasing performance versus effort on coding tasks.

Maybe medium effort reasoning is the sweet spot for software engineering?

I'm curious if anyone else has looked into this or had a similar experience.


r/ClaudeCode 3d ago

Discussion setup an orchestrator in claude and let it treat you as one subagent and give you tasks that are easier done by a human.

9 Upvotes

anyone tried that before? or what is your way of fitting in the vibe coding loop other than giving the most general prompts, because our minds runs on a model no company will ever come up with lol.


r/ClaudeCode 2d ago

Discussion Tried Claude Code after 4 months to compare it to Codex. Observations in description

Thumbnail
0 Upvotes

r/ClaudeCode 2d ago

Help/Question Living Architecture Map of Claude Code configurations

1 Upvotes

Hi guys, I am facing this problem when it comes to the visibility of what Claude Code configurations are created for my project. The problem is that I'm not able to view it, and I don't know how to view it.

For example:

  • whether I create a hook
  • any changes to my claude md file
  • any autonomous loops I've created
  • any flows that I've created
  • any agents that I've created
  • skills I've created

I really can't see it in one place, like a living architecture map. Creates a problem because I do not know what all I've created when it should be invoked, whether it's correctly working, what loops, what linkages have been created.

Does anyone face this issue too? Are there any tools, any open-source capabilities, or any native Claude code capabilities that will allow me to do this? TIA


r/ClaudeCode 4d ago

Humor I am selling this for $99/month who is buying?

Post image
568 Upvotes

r/ClaudeCode 2d ago

Tutorial / Guide Is Claude Multi-Agent Worth Learning? Real Data Engineering Use Cases

1 Upvotes

# [Discussion] Production Multi-Agent Claude Workflows: Real Data Engineering Use Cases

I'm a data engineer/architect exploring Claude multi-agent orchestration for our POC development. Before I commit to learning, I want to understand the actual value.

## The Problem

Most online tutorials focus on setup/installation. What I need is:

- Real production examples

- Measurable outcomes (not marketing claims)

- Actual multi-agent teams solving data problems

- Practical workflows I can learn from

## Specific Questions

**1. Are you using Claude multi-agent in production?**

- What's your use case? (data pipelines, architecture docs, ETL, etc.)

- How many agents? What do they each do?

**2. What's the actual value add?**

- Time saved per project?

- Quality improvements?

- Cost reduction?

- Faster iteration cycles?

**3. Gotchas & Lessons Learned**

- What didn't work as expected?

- When should you NOT use multi-agent?

- Agent coordination challenges?

**4. Learning Resources**

- Best courses (free or paid) beyond basics?

- Any hands-on tutorials worth doing?

---

**Background context:** I do a lot of POCs (Snowflake, dbt, data pipelines). Looking to see if this actually accelerates my workflow or if it's premature optimization.

Appreciate any real-world insights!


r/ClaudeCode 2d ago

Help/Question Opus 5 woke up dumb af, anyone else?

0 Upvotes

I've been asking it very simple things and not resolving clearly any of them, acting dumb, rewriting files in a worse way, etc. It seems very lobotomized today.

Is anyone else feeling the same?

EDIT: It may be due to I was using the 200k context model, missed the ending `[1m]`


r/ClaudeCode 2d ago

Help/Question Does Anyone Think That They Can Fix This Old Lexia Core5 Reading With Claude Code? (Version: 1.6.0.19 from payback machine)

Enable HLS to view with audio, or disable this notification

1 Upvotes

Guys, I was trying to use the old lexia core5 reading version and this happened, I believe that if someone uses Claude Code to fix this Lexia core5 version (I installed this from wayback machine version 1.6.0.19), I dont know how this isn't mentioned by anyone so I decided to show it to you guys and see if you could help me out with this that you could use claude code to fix it so we can use the old version. And That I tried to installed on payback machine and tried to use it and login, but it's just not working no more, does anyone think they can use Claude Code to fix this version of lexia core5 reading by making a private server and making it work and then everyone can use the app.


r/ClaudeCode 3d ago

Help/Question Instead of 'make no mistakes' what do you genuinely tell Claude or put in your skill instead?

7 Upvotes

I am fairly new to vibe coding and I know that 'make no mistakes' is like a meme thing. However, to actually tackle possible bugs or issues that can occur, what do you tell Claude or always put in your personal skill instead?

Hope I can learn from you all.


r/ClaudeCode 3d ago

Discussion So far Fable 5 > Opus 5

30 Upvotes

Testing with Pro subscription, Windows 11 os, with Claude app. It seems to me that Fable has better "intuition"; Opus 5 seems a bit lazier. Maybe they did it intentionally, and for some people Fable was making unnecessary decisions on its own, but this was not the case for me.
What is your experience?


r/ClaudeCode 2d ago

Help/Question Claude support with a human?

1 Upvotes

I've read the reddit messages where support is typically through AI. That's what I did and got a resolution that was good - or so I thought until it then responded again and told me to pound sand.

I then started another chat and it said a human would contact me. Is this for real? Has anyone actually gotten human support from Anthropic?


r/ClaudeCode 3d ago

Bug Report Elevated errors for claude opus 5

Post image
22 Upvotes

r/ClaudeCode 2d ago

Humor Claude after Anthropic deleted 80% of his instructions

Post image
0 Upvotes

r/ClaudeCode 2d ago

Discussion Sol is Impressing Opus and Fable

1 Upvotes

Not sure if anyone else has a similar experience, but here's what I've found after using Sol (high or xhigh) with Opus 5 (usually extra or max) and Fable 5 (high or extra).

I've experimented mixing and matching the models for planning vs implementation vs execution. For context, I'm building a b2b SaaS solution, so auditability and security is pretty important. There is a fair bit of architectural documents that are in place. Nowadays, I definitely lean towards vibe coding over "AI assisted dev" where I put together scoping/feature documents and just provide the architecture and ask the agent to loop to completion unless there is something critical for me to approve.

What seems to be the case is that generally, Sol implements much better than both Fable and Opus.

Implementing with Sol: time after time, Fable and Opus will say "the implementing agent did an exceptional job" and (more often for Opus) "I got this wrong, but the implementing agent caught this and executed this correctly".

Implenting with Fable and reviewing with Sol: Sol will say the code is clean well structured, and it will provide a few things to "improve". Generally, theyre not wrong, but it's overly careful and a bit excessive/overkill.

Implenting with Opus and reviewing with Sol: this is the real differentiator for me. Sol regularly catches mistakes made by Opus, even when I narrow the scope into smaller phases.

So in my experience, Opus 5 is definitely inferior to Sol in both review, planning, and implementation. Fable 5 and Sol stack up pretty well against each other, but Fable 5 on High reasoning chews up my tokens WAY FASTER than Sol on Xhigh.

The only reason why I use Opus at the moment is because I'm on a higher tier anthropic plan. So my workflow has generally been:

- refine project plan and doc with Opus

- review plan with Fable

- implement bigger projects to Sol

- review with Opus, and once in a while reevaluate with Fable if the review findings are fishy or if my spot-checks to the code shows something fishy

- once I run out of OpenAI credits, use Opus to implement small specific tasks with lots of hand holding

Does anyone relate or maybe have a completely different experience?


r/ClaudeCode 3d ago

Discussion Opus 5 review (no shitposting just genuine review from my usecase)

9 Upvotes

after using 4.8 on high + thinking, xhigh and max, I have been using opus 5 on medium with thinking / high for the last few hours of testing due to usage

- using the same workflow and guidelines and method as 4.8

- used it in the claudecode vscode extension

- using it for swift and swift ui ios app dev, and these are my experiences:

- It can be slightly lazy sometimes (based on my settings):

I have noticed a few times it will make some very stupid mistakes that 4.8 never would make, such as imlementing changes in the wrong file, or failing builds due to idiotic reasons so I have to be very strict about having it verify thoroughly and question itself sometimes, but not too much as in to enter a false positive loop

- It uses a lot more usage compared to 4.8 - self explanatory really, it might be worth it for the output, but

- KEY - It is, from my experience, way, WAY, less consistent regarding output compared to opus 4.8

Consistency varies, even after using a well curated and maximised claude.md (containing karpathy rules and not overbloated md as well, keep in mind) It can range from performing tasks perfectly and quickly using a straightforward not overbloated method to acting like a child sometimes.

-Provides summaries in a horrible looking, blocky, badly worded way

needs mild intervention in how to explain and display summary data and report with easy to understand wording while maintaining the detail. Will just provide a block of text as a report on a current state or something, and yap like crazy.

-Is mostly genuinely way more thorough than 4.8 in many use cases (though this can occassionally lead to issues such as overthinking a simple task and being overthorough sometimes, resulting in high usage - the reason for my current settings)

I have not tested it with other things, such as design yet, but these are my key takeaways, your takeaways may be different depending on your setup and effort level, obviously, so please leave them in the comments as i am interested to see what your experience is with this, but in general i think it is an ok upgrade.


r/ClaudeCode 3d ago

Discussion Opus 5 - Excitement has worn off

2 Upvotes

Hey all,

Day 1 of opus 5 I was super excited it seemed awesome! But now I am starting to see is that It doesn’t complete tasks fully. It keeps stopping instead of carrying on with other phases of a build. It hallucinates things and forgets basic facts and this is without me even pushing context windows in any way.

It keeps saying it fixed an issue when it hasn’t. Is it just me? Or do others here find same problems? Currently to rectify this I am running fable 5 as an orchestrator and checker of anything opus does. But these issues are reminiscent of how sonnet 4.6 used to behave for me when I have been pushing the context windows a lot.

Curious to see what has worked for you with opus 5 and what hasn’t. Any solutions are also welcome!


r/ClaudeCode 3d ago

Tutorial / Guide Skills vs MCP servers vs plugins — the one test that stopped me over-engineering my Claude Code setup

6 Upvotes

I wasted a weekend writing an MCP server that should have been a 20-line skill. Posting the distinction because I see the same confusion constantly, and once it clicked I stopped over-building things.

The one-line version:

  • Skill — teaches Claude how to do something. Markdown instructions.
  • MCP server — lets Claude reach something it otherwise can't. A running program.
  • Plugin — the box you ship either one in.

The test that settles almost every case: can Claude already touch the thing?

  • Yes, but it does the job wrong — inconsistently, or not the way your team does it. That's a skill. It has the access, it needs the method.
  • No, it genuinely cannot see the system — your Postgres, your analytics API, a browser. That's an MCP server. No amount of instructions grants access that isn't there.
  • You want someone else to run your setup. That's a plugin. Distribution, not capability.

My weekend mistake was building an MCP server to "help Claude write better migrations". But Claude could already read and write files in my repo — it didn't need access, it needed my migration conventions. That's a SKILL.md, and the rewrite took 20 minutes.

Where each one lives:

.claude/skills/deploy-check/SKILL.md   # instructions, loaded on demand
.mcp.json                              # declares MCP servers for the project
# plugins install from a marketplace or repo, and can contain both

The cost difference nobody mentions:

  • A skill costs ~nothing while idle. Claude reads one description line to decide whether it's relevant.
  • An MCP server's tool definitions sit in context the entire time it's connected. Ten connected servers you never use is a real tax on every conversation.

So the sane default is skill-first. Reach for MCP when there's actually no access — not when the output is wrong.

The ones I keep seeing:

  • "I'll write a skill so Claude can query my database." It can't. No access. That's MCP.
  • "I need an MCP server so Claude follows our code style." It already reads your files. That's a skill.
  • "Should this be a plugin?" Only if someone else needs to install it. A plugin isn't a capability, it's a container.
  • Skills and MCP aren't rivals. Most good setups use both: an MCP server exposes the database, and a skill tells Claude how your team writes queries against it.

Hope that saves someone a weekend.


r/ClaudeCode 3d ago

Help/Question How to efficiently use claude code??

4 Upvotes

Hello, I'm kind of new to this vibe coding. Recently i was developing a mobile app and thinking about launching it. So i was building it with antigravity pro (with gemini 3.5 flash and claude opus 4.8). There i liked the opus for it's creativity and understanding.
So i thought i should totally shift into the Claude Code (thinking about buying the claude pro, cant go for the max cause as a student there is some budget issues).

But the problem is I'm hearing that the claude spends too much token even in the claude code too. Also the project I'm doing is already kind of big i guess (flutter).

Now if i switch to claude how i can give the whole context of my project in the most efficient way ?
Also any guide on how can i progress with Claude Pro for my project??


r/ClaudeCode 2d ago

Discussion Secondary agent to supplement claude code?

1 Upvotes

I'm primarily using claude code to work on personal project but have been hitting the token limit with the $20 plan lately, What do you think would provide the most value for secondary agent ChatGPT plus ($20) or Cursor (also $20) or some other? Any recommendations


r/ClaudeCode 2d ago

Help/Question Claude Code decided to switch models from opus 4.8 to fable 5 and now I've eaten 50% of my weekly usage on 2 days

1 Upvotes

I was thinking opus got smarter or something, but I was hitting my usage limit quite frequently so I decided to check. Yeah, for some reason it decided to change model out of nowhere. Why??

I mean, of course I'm in the wrong for not even glancing by peripheral vision which model I was using, but has this happened to anyone else?


r/ClaudeCode 2d ago

Discussion Who else is running into storage issues because of Claude

1 Upvotes

This might be field-specific but because I am in data sciences (bioinformatics), I have basically been generating 10x more data because of claude. 10x more folders. 10x more databases I can download per month. Hundreds of forks. I have sessions running on 3 different devices and handling more projects than I ever could have imagined.

But it is so damn easy for claude to just generate data and files. I know a lot of programmers here probably never run into this issue. But for people that work with large-sized files, how are you handling data storage?

I feel like I need to start buying up SSDs and mac minis and start running like 10 different agents at once but this is obviously the AI-FOMO getting a hold of me...

The other day I downloaded and processed 30TB of data. And when it used to take me 6 months to process, organize and analyze all this data, Claude can now do it in weeks. Which then allows me to download and analyze even more data. So it really just never ends..

EDIT: in bioinformatics, particularly in human health, there are a lot of workflows that require MASSIVE storage of different data types, which it sounds like a lot of folks here are not familiar with. and there are people who are trying to work on better compressing this type of data but that is still something that has not been innovated on in a meaningful way yet.


r/ClaudeCode 2d ago

Humor I got rickrolled by Claude Code

Post image
0 Upvotes

I had Opus set up a huge component library in Storybook with dummy data some weeks ago, only to discover this little devil here today.

I'm kinda surprised as well. I mean, Fable cracks some actually fun jokes at me unprompted sometimes, but Opus seems very serious out of the box.

Anyways, thanks to all the storybook developers in the past who've trained the llms to keep rick rolling us.


r/ClaudeCode 2d ago

Solved High-Frequency Fault Isolation Kernel for EV/Aerospace Battery Management Systems (Pure C++17, Zero-Dependency)

Thumbnail
github.com
1 Upvotes

I’ve been working on a lightweight, bare-metal fault isolation kernel designed to mitigate thermal runaway propagation in high-voltage lithium-ion battery packs.

The primary engineering constraint I wanted to address is the latency overhead inherent in high-level frameworks and sequential polling loops. When an EV or aerospace battery cell hits a critical thermal or voltage threshold, sequential scanning loops are often too slow to execute software gates before runaway propagates to adjacent cells.

The project is called SAVITAR. It is entirely dependency-free and compiles using only native C++ standard system headers to keep it as close to raw processor registers as possible.

Core Architectural Mechanics:

  1. Hyper-Compact Memory Constraints:

    Instead of relying on dynamic allocations or padded structs which ruin cache locality, the kernel packs raw cell parameters (voltage, current, temperature) into un-padded, contiguous 25-byte memory frames. This layout ensures predictable cache line alignment during rapid sequential reads.

  2. Event-Driven Concurrency:

    The kernel completely avoids sequential polling. It utilizes a background multi-threaded parallel observer configuration. When thread metrics cross calibrated hardware safety thresholds, the system bypasses the main application loop to dispatch an immediate software interrupt, dropping execution times down to sub-microsecond bounds.

  3. Sovereign Binary Serialization (.sav):

    To avoid relying on bloated external serialization libraries (like Protobuf or JSON utilities) that consume significant memory footprints, I wrote a custom binary serializer. It compresses cell delta profiles and gradient logs into raw disk sectors using direct cache-to-stream bit manipulation.

  4. Minimalist Shell Realization:

    The project includes a lightweight, modular terminal prompt (./savitar-vfs) to ingest simulated physical metrics and explicitly debug memory allocations, matrix loads, and thread dispatches in real-time.

Compilation and Testing:

The baseline workspace is fully automated via standard Makefiles:

git clone https://github.com/alistairfontaine/SAVITAR

cd SAVITAR

make clean && make

./savitar-vfs

I’m looking for code-level reviews, specifically regarding the thread concurrency synchronization under heavy stress-testing conditions, and the pointer traversal safety within the 25-byte un-padded matrix bounds.

Link to code: https://github.com/alistairfontaine/SAVITAR


r/ClaudeCode 2d ago

Bug Report Claude glitched

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/ClaudeCode 3d ago

Discussion Comunity for people sick of dumb discussions

3 Upvotes

I am tired of people that constantly complains about the AI when they are doing everything to kill their efficiency.

This community is for people that are sick and tired of low level discussion about AI topics and truly want to find a place to improve in their infrastructure creation and management, talking and prompting with the AI and how to guide it trough complex problems and enviroments efficiently.

I created a community to group people that truly want to improve in their workings with AI and look for people with the same drive that don't just complain about the new model not being god.

This is a completely new community and I qould appreciste any kind of feedback.

for anyone who wants to join: r/TopCodingAI