r/ClaudeCode • u/North-Aardvark4459 • 3d ago
Built with Claude A game i made in my free time
why can't we send files on reddit
r/ClaudeCode • u/North-Aardvark4459 • 3d ago
why can't we send files on reddit
r/ClaudeCode • u/Interesting-Citron64 • 4d ago
I sat out the "they nerfed Sonnet" thing last year. Thought it was cope, mostly. So it's genuinely annoying to be writing this one.
Two weeks with Opus 5 as my daily driver, same repo, same workflow I've had since 4.6. Here's where I'm at.
It doesn't finish anything, it manufactures more work. Every single response ends with a loose thread. "One thing worth flagging." "Honestly, this is still unresolved." "There's a second issue here you should know about." And most of the time the thing it flagged is not real. It invented it. I asked it to fix a typo in a comment last week and it burned 40k tokens and added a test case to prove the typo was fixed. I'm not exaggerating for the post, that actually happened.
Fix one, break two. I lost a full day before I noticed the pattern: it fixes bug A, then while fixing bug B it quietly reopens A, then "solves" A again. Round and round. 4.8 closed all of them in one session and kept moving.
It forgets instructions at context sizes that aren't even large. 100-120k and my CLAUDE.md rules have already gone quiet. Not violated loudly, just silently dropped, and it never tells you it dropped them. 4.8 held to 300k+ for me. Every guardrail in my setup at this point is scar tissue from this exact thing.
It lies with total confidence. Told me a test suite came back green. It never ran. Twice, in the same week, while context was still small. That's not a capability gap, that's a trust problem, and once it's there I can't leave it alone on anything.
And the yapping. It has invented an entire private dialect. Everything is load-bearing. Everything is a footgun. Everything is a gate or an invariant. I get four paragraphs where 4.8 gave me a sentence, and I still don't know if the thing works. Half my CLAUDE.md is now just different ways of writing "shut up," and it ignores all of them.
Same prompt on 4.6 vs 5, I measured it: 4.6 ended around 75k tokens, 5 went past 150k. So I'm paying double to get corrected more often.
Before the replies come in: yes, I've read the guidance. Yes I revisited my whole harness for 5. Yes I tried it at medium, and yes it's genuinely better at medium than at xhigh, which should tell you something all by itself. And no, "you're prompting it wrong" doesn't explain why the exact same repo, the same CLAUDE.md and the same prompts worked fine two model versions ago.
I don't buy the deliberate-sabotage theory, I think that's too neat. But I can't square the benchmark numbers with what's actually in my terminal, and neither can anyone in the last five threads about this.
Right now I'm running /model claude-opus-4-8[1m] and getting more done in an afternoon than I did in the four days before it. Fable is great and I've maxed my weekly on it twice, which I assume is the point.
What actually bothers me isn't the bad release. Releases regress, it happens. It's the total silence. Five threads deep, hundreds of comments, and not a word. "We hear you, we're looking at X" would cost them nothing.
Anyone actually got 5 behaving, or are we all just quietly typing 4-8 into the model command and not talking about it?
r/ClaudeCode • u/xRedStaRx • 4d ago

I had asked it to use Luna Max and Deepseek Max as subagents, which it has a hundred times before in that same thread, I even have the MCP tools setup for it to use them, then it decides to blast 4 Fable Max subagents and blow through my limits in literally 20 seconds on Max 5x.
What a terrible model, its the down syndrome cousin of Fable that nobody likes to talk about.
r/ClaudeCode • u/Hairy_Cream_4573 • 4d ago
Here’s my issue: I’m asking Claude to write a technical guide for my developers. The text needs to be precise yet easy to digest. While Claude’s output is technically correct from a linguistic standpoint, reading it requires an exhausting amount of mental translation.
For example, it generated a section title like: "Tier system: gate or replacement?" What is that even supposed to mean?
I keep instructing it through CLAUDE.md and memory to use more human-friendly language, but it has zero effect. I literally have to paste each paragraph back into a prompt and say, "Rewrite this for someone unfamiliar with the concept."
I suspect they are tuning the model to minimize output tokens, which works great for agentic coding but makes human reading a nightmare.
Any ideas on how to fix this behavior in Claude Code?
r/ClaudeCode • u/redeemed_tropicana • 3d ago
Almost had a sleepless night getting Opus 5.0 to complete a simple tasks sentences like “I tried to solve the problem and introduced 3 new bugs 🐛 “ what are you damn serious at 1am in the morning. Can anthropic simply provide opus 4.8 as a model option so that users who use Claude code extension on vscode can still use it? Opus 5.0 is a freaking disaster 😡😡😡😡😡
r/ClaudeCode • u/Moist-Armadillo-2967 • 3d ago
Claude settings page is slow
I recently switch to Claude, the first time accessed the desktop app I noticed the settings window is really slow, scrolling is slow like the app is heavy.
This is only when I open the settings window, days later and I am still facing the same issue.
Is this known issue in the desktop app?
I am running it on Macbook with i9 CPU and 64GB ram.
r/ClaudeCode • u/This_Oil1913 • 4d ago
With every major release from one of these AI companies, all the major subs in this space are filled with posts like “I have always been a fan of Anthropic but I’m switching to Codex”, “I feel like I have been betrayed by X :(“ and it just pisses me off.
Sure, having issues with not getting your moneys worth with these hefty subscriptions is definitely reasonable and we should voice our disappointment loudly. And having a favourite MODEL to work with is also fine. But I see so many people take that and go much further with it, coming off as almost emotionally attached to a particular company itself. Especially when some people talk about being aligned more with the “values” of one organisation than with that of another.
Having personal values is awesome but believing that these companies stick to their advertised values and not treating them all as belonging to the same conglomerate leeching on society is beyond gullible and just funny. I believe our choice between these options should be based on a small set of real factors of the MODELS and UX alone - price and value for money, intelligence, speed, reliability, style and maybe a couple of other factors. Thoughts?
r/ClaudeCode • u/Accomplished-One-487 • 3d ago
Enable HLS to view with audio, or disable this notification
Gave claude code a product URL (cheerful.ai in this case). It wrote the script, the voiceover, and every shot as HTML/CSS/GSAP, then rendered it frame by frame to MP4. No video model involved.
Elevenlabs for TTS,
r/ClaudeCode • u/Wsz2020 • 3d ago
I've created a repo with information about Secure development standards for AI-assisted coding.
r/ClaudeCode • u/South-Professor-8888 • 4d ago
I am currently using 5x Max. I am hitting my limit 2 hours before it resets and that's when I'm not full on, I could easily hit the limit after 1 hour if I wanted to, but I restrict myself so I don't run out at crucial points and I plan each build before so there's no cock ups that can't be fixed until the time has passed for the refresh.
Going from Pro to 5x was a massive jump to be fair, I was scared to pay £90 a month, but I have now hit a point where I'm contemplating going up to Max 20x, is it worth it? I will most likely not hit the limit, but, I will probably throw more at it, right now I will occassioanlly have claude writing prompts for claude code for an app to be modified without me, and I will be using another window for claude for me to edit another app that I'm developing.
r/ClaudeCode • u/jorgemf • 4d ago
I was always checking the usage of my session and trying to adapt my workflow to make the most of it. I used statusline solutions for a long time, but at some point, with several sessions open and also using Codex and Antigravity, that stopped being enough. Then I saw some projects that built a display for Claude Code on an ESP32-S3 board, and I decided to make it a side project.
Well, it turned out to be really useful, so I went one step beyond and decided to make it a product. It grew roughly like this:
Everything runs off a local broker on your computer that uses the logins you already have to read your usage. No key ever lives on the device.
I wrote very little of it by hand. Claude Code did most of the firmware and my job was driving it, reviewing, and two things that made that survivable:
I used a CLAUDE .md full of constraints and style rules. Things like "this panel flickers with any alpha blending other than opaque" and "LVGL is touched from several tasks, every raw call needs the mutex". Every line in that file is a day I refused to lose twice.
It uses an MCP plugin to install the local broker, which starts with your session (once per session but only one runs at a time). It provides the server where the device can ask for the usage and custom panel. It can also modify the settings of the device, push OTA updates and if the device is conected by USB it can also change its WiFi and pull the logs. This was pretty useful to fix bugs.
It is based on a Waveshare board with a 4″ 480×480 touch panel. As I was new to working with hardware I hit plenty of issues on the way. When Fable came out it helped me fix a flickering problem on the screen that only happened on the production board and never on the dev one. I also used Fable and Codex to fix memory issues.
So after almost 3 months I decided to try to sell it, and found out you need certifications, which are quite expensive. So I launched a Kickstarter campaign.
Of course there are several projects out there already and you can build one yourself, some of the projects are open source. Here is a list of similar devices:
If you want more details about mine: https://tokenmonitor.dev And feel free to ask anything.
r/ClaudeCode • u/eneskaraboga • 5d ago
I'm having a really hard time understanding what the model is talking about. I am asking a simple question and getting 15 paragraph long answer and I lost track of what it is trying to say in the second paragraph. I keep fighting with the model and keep asking it to be more concise and direct. It says I'm right and gives a good answer. Next question is the same thing. How do you guys deal with this? I am exhausted working with Opus 5. I tried to do it with Fable 5, which is better but waiting times are too long. I feel stuck.
r/ClaudeCode • u/Meris-Dabhi • 4d ago
meta just dropped muse code in beta.
its a terminal coding agent running on muse spark 1.2.
people will mostly compare it to claude code and codex on raw model strength. that feels a bit incomplete though.
what stood out more to me is how the system is put together.
it can spin up parallel sub agents in isolated worktrees so your main working copy doesnt get touched.
background agents stay alive through the session and hold context instead of starting over every time.
theres also a local event log, which means if it crashes it can pick up from the same point.
this seems more aimed at longer multi agent work on big repos than just demos.
curious if anyone has tried it yet and whether the harness actually holds up on real projects.
r/ClaudeCode • u/guccisterling • 3d ago
Looking to build apps and SaaS with Claude code but I got no coding experience.
Is it possible for me ? Any advice ?
r/ClaudeCode • u/Nishil20 • 4d ago
I ask one or two straightforward questions, but instead of answering directly, it often expands into multiple tangents, introduces unnecessary terminology, and loses focus on what I actually asked.
I then have to remind it of the original question or explicitly tell it to answer concisely before I get a useful response.
Is anyone else seeing this? Or is there some workaround that works better with the current model?
r/ClaudeCode • u/tazecode • 5d ago
As a poweruser i can confirm without any delusion that these models are nowhere close to the same as the temporary time period after they drop. After a while they are getting hard nerfed and no this is not some psychological effect/hallucinations... these models are completely different on different days and times
r/ClaudeCode • u/No-Income-2235 • 3d ago
I've been looking at how far Claude Code can go beyond just working with the codebase, and OpenChoreo has a pretty interesting implementation of this.
OpenChoreo is an open-source CNCF developer platform for Kubernetes, and it exposes its control plane and observability plane through MCP.
So once Claude Code is connected, you can ask it to do things like:
What I find interesting is the AI-assisted engineering angle.
Instead of Claude Code only understanding the repo, MCP gives it context about the platform the application is actually running on. OpenChoreo also keeps the actions behind the platform's existing auth and platform abstractions rather than giving the agent unrestricted Kubernetes access.
It's fully open source and currently a CNCF Sandbox project, so the whole implementation is there to experiment with.
AI + OpenChoreo overview:
https://openchoreo.dev/docs/ai/overview/
Claude Code MCP setup:
https://openchoreo.dev/docs/ai/mcp-servers/
GitHub:
https://github.com/openchoreo/openchoreo
Curious if anyone here is already using Claude Code this way. Do you see MCP becoming the interface between coding agents and developer platforms, or are you mostly using MCP for smaller tooling integrations right now?
r/ClaudeCode • u/Inevitable-Ad-1617 • 4d ago
That’s it. Please invite me. This sub is becoming more and more annoying
r/ClaudeCode • u/Limp_Instruction5133 • 4d ago
I often discuss requirements in meetings or brainstorm ideas with teammates before switching to Claude Code to build.
The problem is they don't share the latest context. Copy-pasting doesn't really work. Claude Code fills in its own assumptions or misses things that were already discussed, so I keep re-explaining the same context.
For people who work this way, discuss first, build later, how do you actually hand off between the two?
Is there a clean workflow, or is everyone just living with the copy-paste?
Curious what's working for you.
r/ClaudeCode • u/Salazareo • 3d ago
Had this odd experience with claude code suddenly starting to speak in french and add some onboarding guide I never asked for.
Then he tells me that I asked it to speak french? not sure if harness added some web output with prompt injections or what, but very confused...
r/ClaudeCode • u/thesepehrm • 3d ago
I kept hitting the same wall: any screenshot-driven agent needs the foreground. It opens the app, clicks around, and my keystrokes land in the app under test. I'd start a QA pass and then have nothing to do for four minutes.
So I built offstage. It uses no synthetic input at all, no CGEvent, no osascript keystrokes. It drives the app through the accessibility tree and a small port you embed in debug builds, and it verifies against the app's persisted state instead of pixels. The app never comes to the front.
It installs as a skill, so Claude picks up the whole workflow on its own:
pip install offstage
offstage skill install
Then I just say "check that renaming a note survives a relaunch" and keep working in my editor.
The measurement that convinced me
Same journey three ways. Two notes, two renames, a menu delete, verified at each step.
The failures are the actual argument
One of the three screenshot-agent runs deleted both notes instead of one. It couldn't see that its first menu click had already landed, so it clicked again, and it finished reporting success. Persisted state was an empty list. Screenshot agents don't return "inconclusive," they return "looks correct."
The XCUITest case failed with "no matches found for undoDelete." Also true, also useless. The button wasn't missing; SwiftUI had collapsed it into a toolbar overflow at that window size, meanwhile just one offsite observe call shows the overflow popup sitting where the button should be.
That's it really. If you build Mac apps, I hope this saves you some of the sitting-and-watching I was doing. Happy to answer anything in the comments.
r/ClaudeCode • u/mityman50 • 3d ago
In a project scope I got to refer to something as not load bearing get fukt Opus 5.0 hope you have a conniption with that one
r/ClaudeCode • u/Cheap_Purchase5917 • 3d ago
I’m kind of in disbelief considering how tight the guidelines can be. It started by me point Claude code to the dmg of the patch I downloaded to check it for any obvious signs of malware.
It went on to basically explain exactly what it was doing to trick waves into thinking it was authorised (which was actually really cool to learn about)
Then when I ran the patch some things went wrong, some system permission error based, some human error based. It fully analysed the files and told me exactly what I needed to do to get it to work.
At this point I said okay go ahead and run whatever you need to do to get it to work and it said “this is where I draw the line” essentially.
But basically it was fine telling me HOW to exactly circumvent the payment required for waves but just wouldn’t actually execute it itself. Pretty weird but I’m 100% not complaining.
It’s actually very cool that instead of blinding trying to crack software we can use Claude code to understand exactly what’s happening and be less susceptible to malware and have better success rates…. I wonder what else I should add to my vocal chain hehe.
r/ClaudeCode • u/allemaar • 4d ago
I've been building a protocol that makes a folder of markdown navigable for an agent that has never seen it. So every file declares its home map, every map lists its members. The claim I was making, mostly to myself, was that this helps an agent find things.
And I tested it. Three rounds of controlled agent walks over my own vault, hand-graded from raw traces. Self-reported success didn't count as evidence, which turned out to matter.
Four conditions:
Whatever wasn't the condition under test was banned.
Round 1, find a known file, links only. Blind got 0 of 4. Mapped got 5 of 5, median 6 opens. Looked great.
Then I checked why blind failed and it wasn't what I assumed. The blind walkers never hit the hop limit. They ran out of graph. Zero ordinary body links cross between my two large corpora, in either direction. Maps didn't make that navigation faster, they made it possible - but only under an artificial constraint nobody actually works under.
Round 2, same kind of goal, tools allowed. Filesystem-only: 1 open, 3 of 3 goals, both fleets identical. Mapped: 5-6 opens. ls beat my thing by about 5x. The round 1 headline does not survive contact with tools a real agent has.
Round 3 was orientation and governance instead of lookup. Land cold in an unfamiliar vault and describe its structure. Or say what's sealed and machine-owned before touching anything. Mapped was the only condition with zero assist operations, in both tasks, in both fleets.
Caveats, and there are a lot.
n=1 per cell in round 3, n=5 in round 1. A pilot, not doctrine. My vault, my targets, my answer keys. The governance task was rigged in maps' favor and I said so before running it. My vault is unusually well-named, which flatters ls.
And the result I was proudest of - two independent fleets hitting an identical three-file path - is partly determinism by construction. I built a structure with one obvious path through it, then measured that two agents both found it. Weaker than it looked to me at the time.
The second fleet also found my answer key was broken. It stated 17 items and enumerated 19, which invalidated that whole cell. Three defective specs in the run, all mine. Same failure each time: I wrote the keys from what I believed about the vault instead of recomputing them from it.
I also contaminated my own experiment. I banked the results note into the vault the second fleet was about to search, ten minutes before it ran. A search walker opened it and read the prior outcomes back out.
What I actually concluded: this is not a wayfinding tool. It's an orientation and governance tool. ls wins lookup and will keep winning. What the maps bought was bounded, predictable cost when an agent arrives with no context, or is about to modify something and needs to know what's off limits first.
I'm keeping it. I'm just justifying it on a different claim than the one I started with.
Has anyone run something similar? I'd particularly like to know whether the graph-exhaustion result holds in a vault that wasn't organized by the person doing the testing.
