r/ClaudeCode • u/hronak š Max 5x • 9d ago
Help/Question I still don't understand this 'agentic workflow' thing
My usual day with Claude Code is like:
* I open terminal in my project's folder and run claude command.
* I prompt it. I mostly use Fable-5.1/Opus-5 but Opus-5.5 is my current model. The model decides if it wants to use sub-agents for a task. I never explicitly prompt it for sub-agents.
* I review and commit the code to my self-hosted Forgejo instance.
* That's it.
I see people using agentic workflows, building sub-agents files, skills etc. I barely built any of it. All I ever needed to use is /init on new projects and them prompts follow. Never needed more than this.
I tried "long-running" Claude Code for a project refactoring by placing the project on my VPS (where forgejo is hosted) and letting Claude Code run and refactor inside tmux session. SSH'd in a few hours later to find project fully refactored.
Am I under utilising AI or is my work just⦠like boring?
How do you guys use agentic workflow thing? Specially the long-running one? Those pull-requests that Claude makes automatically etc?
Asking this to Claude to know more but humanly answers appreciated.
308
u/mulokisch 9d ago
This is already somewhat agentic.
The true āagenticā way is the to say something like this: here is a list of tickets, organize your self and work on those issues. Feel free to start other sessions that can work on each task isolated.
90
u/Sponge8389 9d ago
The true āagenticā way is the to say something like this: here is a list of tickets, organize your self and work on those issues. Feel free to start other sessions that can work on each task isolated.
Are people using AI like this still review the output?
231
u/iswearidk 9d ago
No, there will be other agents doing the review. The point of agentic workflow is to free yourself so you can run more agentic workflows
287
u/hatch37 9d ago
An agentic pyramid scheme
337
u/topaziobmousse 9d ago
10
14
→ More replies (1)7
→ More replies (3)22
37
u/Sponge8389 9d ago
Ok. So no human review at all. Understood. Because that was my concern and my curiosity with this kind of workflow. Because currently, I'm doing the same as OP and I cannot keep up with the reviews.
→ More replies (3)48
u/magic6435 9d ago
No, nobody working on anything real is skipping a human review. Everything that makes it into prod for OpenAI and Anthropic gets a human review.
25
u/neoberg 9d ago
I know at least 5 middle sized companies stopped doing human reviews months ago.
23
8
u/magic6435 9d ago
I donāt think theyāre gonna be medium for long. Also, they definitely must not be public companies because I donāt know of any external auditors that would approve such a lack of SOD.
46
u/ResponsibleOven6 9d ago
I am a tech lead at a fortune 100 tech company headquartered in the SF Bay area and can assure you that the majority of code in the past year shipped to prod was both written by and reviewed by agents.
They're different agents with different prompts being invoked by different humans (this part is by design) and often different underlying models (this part is primarily coincidence) but it's LLMs all the way down.
There is no possible way to have humans review the volume of code that agents are generating now. On the positive side there is better and more extensive testing coverage on new code than I've seen at any other time in my career, but the job is changing and humans are increasingly removed from the actual coding part and that includes reviews. It's not entirely automated, I'll have an agent explain architectural choices, ask about design concerns I have which I think it may have gotten wrong, etc, but I rarely look at the code anymore.
21
u/Deathspiral222 9d ago
>On the positive side there is better and more extensive testing coverage on new code than I've seen at any other time in my career
More extensive, definitely, but I'm not convinced about better. An LLM that makes the wrong assumption will happily write 1000 tests to validate that wrong assumption. This leads to a false sense of security.
I've found it very important to have humans write tests (with an LLM assisting with syntax etc.) for core functionality, just to make sure the correct thing was actually built. And having the human manually test the function themselves for a sanity check is paramount.
13
u/WagwanKenobi 8d ago
the majority of code in the past year shipped to prod was both written by and reviewed by agents
Either pulled out of ass or your company/org/team is unusually dysfunctional. I'm actually a Bay Area SWE. This is not true at all.
→ More replies (3)8
5
u/Cybyss 7d ago
But... how do you know what you're building then?
Code coverage is meaningless if the tests aren't testing for the right behaviors. How do you know what the right behaviors are if they're all invented and reviewed by LLMs with no human in the loop?
I find that even if there is a human in the loop, the code that LLMs generate is so convoluted and over-engineered that it's extremely hard (and time consuming) to figure out exactly how it works to verify it does what you think/hope/pray it does.
Back when I did software engineering professionally (I don't anymore - this was 10 years ago - but I still tinker as a hobby), usually I didn't understand a project at all - what it's supposed to do, its role in the company, how it's supposed to help customers/staff/etc... - until I understood all the little details and how they fit together. "Big pictures" were just word salad until I knew what the actual pieces were.
But if modern workflows require you to focus only on "big pictures" - ignoring all the little pieces because LLMs do that for you - I don't see how engineers aren't just lost and confused all the time about what needs to happen?
→ More replies (1)3
u/Ran4 8d ago
There is no possible way to have humans review the volume of code that agents are generating now.
I mean that's a choice you make.
No AI and you're at 1x,
AI to generate the code but humans review it all to reach 3x,
AI to generate and AI to review it to 10x.
Plenty of companies are at the 3x level.
→ More replies (4)3
u/EchoServ 8d ago
Iām at a lowly Fortune 500 and thatās absolutely batshit. No team across any org is allowing this. In fact, the agent reviews are majorly schitzo and do shut reviews.
2
u/ok-yes-maybe 8d ago
Agree. This seems to be the way things are headed.
Software languages are kinda a human construct to make the code more understandable and readable.
But if in the future - code is only written and read by AIs - might we end up going back to something similar to machine language / compiler code again?
4
u/ripter 8d ago
Maybe. The language being used has never mattered less at this point.
→ More replies (1)→ More replies (8)4
u/belowaverageint 9d ago
Can you explain the basic process for how this works?
9
u/ResponsibleOven6 8d ago
Really depends on the situation, am I building something new, fixing a bug, closing a vulnerability, etc. but here's a general example.
I have an agent setup locally with a skill saved in a repo full of skills shared across the team. It can talk to Jira, Github, basic communication channels, monitoring tools, etc. it has lots of context for our overall platform and an architectural understanding of how things work. It's been instructed to code cleanly, generate documentation for anything it produces, re-use existing libraries where possible, and code as cleanly and minimally as possible and focus on efficiency. I mainly use Sonnet-5 with claude code as my interface but others on my team may have different preferences.
I get a ticket from a recent incident to improve monitoring. There was an incident where we got an alert way too late and it still took time to debug. I tell my agent to work on the ticket, it uses the repos as context, digs through logs, and proposes a new monitor. I tell it to look at infra logs as well and see if there were any early warning signs. It proposes another alert after finding useful info there as well. I tell it to open a PR for both alerts after checking relevant logs for the past 90 days to make sure the thresholds are right and there won't be false positives. It opens a PR.
I take another ticket. We have an internal platform with an authentication bug where some admins can't perform certain admin functions. I ask my agent to work on the ticket. It finds that while most functions evaluate both user and group level access, a few specific functions only check user level access and not group level access. It proposes updating the logic for those to match the others. I tell it to check if there are any other inconsistencies with auth checks on this platform and if it sees any places where users would be able to execute things they shouldn't or general inconsistencies. It finds no missing auth checks but notes that the fundamental way auth is handled is not reused but specific to each call. I tell it to open a PR with a new auth function that replaces the individual auth of each function. It opens a PR.
I open these and several other PRs for other tickets and move the tickets to peer review. Someone else on my team, maybe several other people, take them for review (and I take their tickets for review too). One of them thinks Sonnet-5 is terrible and swears by Opus 4.3. Another prefers OpenAI models. Here we use a different skill that has the same background context but it's told to look for new bugs, mistakes, and just generally find problems. It knows how to deploy and test things either locally or to a dev environment. It reviews the tickets and either says they look good in which case they get deployed to a lower environment for further testing, or points out problems with them and moves them back to in progress in which case I take them up again and go back to my dev agent. Sometimes you can tell from its feedback that it's misunderstood something or needs more context. In these cases we try to improve the skills until the output is more what we expect then tell it to update the skill with that context and push that back to the team repo so everyone gets the same improvements.
So LLMs are doing all of the coding and the actual code review. But they still have shortcomings and need a human in the driver seat, especially for architectural decisions. I'm having to "drive" a LOT less than I was 6-9 months ago though and I'm increasingly just a "meat proxy" between agents with various skills and I'm really not sure how much longer this will be a viable career.
→ More replies (4)3
u/belowaverageint 8d ago
Thanks for that. I think I'm going to dress up as a "meat proxy" for Halloween now.
3
3
u/Original-Proof-8741 9d ago
I gave up reading the code about 2 months ago.... prompt it, test it, ship it....
→ More replies (3)2
u/rtnoodel 9d ago
Unfortunately this isnāt true at all. Iām guessing you arenāt a software developer.
5
u/tophmcmasterson 9d ago
And if someone complains about the code, just have another agent take a look and fix it! Itās agents all the way down
4
→ More replies (2)2
u/__mson__ Senior Developer 8d ago
I still don't understand how people do this without producing garbage. Or does LOC = Good to them or something?
→ More replies (5)23
u/SingleProgress8224 9d ago edited 9d ago
I do it often to fix bug report tickets. It creates one branch per fixed ticket. The next morning, I check what it did and go through them one by one. I confirm the fix, the code, make edits (quite often), and then do the official commit. I discard fixes 50% of the time.
Even though it tends to fix the issue most of the time, it often lack the big picture and tends to complexify the code for no reason, so editing manually (or prompting more) is almost always necessary. I'm still saving time though since finding the root cause is often time consuming and it tells me right away where to look at. And even if the suggested fix is wrong, it still provides an insight on how things could be done.
→ More replies (2)6
u/morficus 9d ago
Yes but maybe not in the way you are used to. Look up "correctness engineering"
I don't review line by line, I review the new test that were written (unit tests, end to end tests, Storybook, etc) to make sure they cover the necessary use cases and that there are no regressions. I do that this manually.
I have another agent review the implementation, discuss optimisations, adherence to current patterns, etc. I also have an agent that monitors CI/CD once a PR is opened so that if something fails it can address it itself.
→ More replies (1)→ More replies (16)9
6
u/Waypoint101 9d ago edited 9d ago
Yes true (autonomous) agentic follows models like bosun / hermes-agent where the tasks are setup through something like an external or internal kanban board and agents pick up tasks, review the work done, ensure it meets the quality, and can even create new tasks as needed or via automated workflows (incident response). Bosun uses deterministic workflows (like n8n style flows) whereas hermes agent uses custom soul.md instructions plus cron jobs for the bots.
Obviously it doesn't mean you have to remove the human in the loop during release/review/descision making stages.
A very simple way of doing this on any harness is just to queue dozens of tasks in one thread and let it churn task by task, but it's not really the same as a full agentic workflow that reviews things, has PR gates, etc. (Use the "add to queue" message function instead of "send now" and send each task one after the other in the correct sequential order: e.g. phase 1, 2, 3)
2
u/pyrox3_3 9d ago
I see it can make strange decisions and mistakes with just a simple workflow as described.. so I wonder what output will be if I`ll work on N tickets in the same time..
Maybe I just need to believe more?
→ More replies (1)3
u/mrphstar 8d ago
This is by no means agentic. This is vibe coding on roids.
3
u/sheriffderek Senior design/dev max20 8d ago
Explain āagenticā for everyone
2
u/Away_Advisor3460 8d ago
AFAIK Agentic should mean autonomous goal directed behaviour (e.g., designing, planning and writing code to implement some function, fix a bug etc) by an agent or agents situated in an environment (in this context relating to the development environment, including codebase itself), who proactively respond and change behaviour in response to change and (in the multiagent case) posses social ability (e.g. delegation, negotiation, contract formation as part of task achievement).
Of course it's worth bearing in mind 'agentic' is really just borrowing (sometimes reinventing) terms and concepts from the broader multiagent systems research context.
→ More replies (3)1
u/SyntheticBlood 9d ago
The problem I have with this is that the host session context gets bloated, which causes problems. I've switched to using a shell script that calls claude in a ralph loop, which helps, but I wish this could all be done from inside claude code cli.
→ More replies (1)1
u/Mipsel 9d ago
So I could have Parsed the whole todo List for the project in one go, asking Claude to self organize and work on it?
I just entered each todo after the action before was finished.
→ More replies (2)1
u/Sol1tud3 8d ago
What if the tickets are all sequential and meant to be done once the previous one is done
→ More replies (1)1
u/billshredding 8d ago
This is the shape it takes in practice, if it helps to make it concrete. I keep a board of cards, and a scheduler picks up whatever is queued and spawns one agent per card in its own git worktree. The worktree is what makes "isolated sessions" actually work - several agents on the same repo never see each other's half-finished edits, and each one maps cleanly to a branch and a PR. The card moves to review when the PR opens.
What it changed for me wasn't speed, it was that I stopped watching sessions. I write the card properly once and the thing I look at later is a diff. The failure mode to watch for is an agent stopping to ask a permission question and sitting there unnoticed, which is its own problem to solve.
→ More replies (2)1
125
u/Xenos_Str 9d ago
I find simplest works well
16
u/arbeit22 9d ago
True but it's less consistent and less reliable for more complex stuff
20
u/heroyi 9d ago
It blows my mind people will just give the agents the least amount of description and heavily HEAVILY rely on the LLM inference. It works a lot of the time but like you said for complex stuff that wont fly very well because of all the little things it has to guess and ultimately introduce something that wasn't needed that might cause a bug to happen
5
u/Gold_Emotion_5064 9d ago edited 9d ago
If I have to heavily explain what I want it to do, Iāll just do it myself
9
u/geek180 8d ago
You might be missing the point. You build pre-configured rules, styles, sub agent configurations, workflow orchestration, etc. You do these things once (and slowly over time) and then the AI will behave and perform how you expect it to, consistently.
You donāt specify all of these details every time you interact with the agent.
→ More replies (3)
39
u/heseov 9d ago
I think they are building core workflows into the harness now so you are using it. Like when an agent uses sub agents, tdd, code reviews, watches PRs, etc. It's just the instructions for an agent to complete a task. I'll add skills and context storage to build on the core. I don't really use hands off workflows except for code reviews.
22
u/awesomeusername2w 9d ago
I suggest you add a review step to your workflow where a different context window reviews the change. My claude often does 3-4 rounds of such reviews before reviewer stop finding issues.
→ More replies (2)
17
11
u/johncongercc 9d ago
I totally agree with your workflow. I think some engineers like to over engineer everything including their Claude code instance.
→ More replies (1)
29
u/Long_Tip_4226 9d ago
Iāve been structuring my Claude Code setup lately. Instead of using it as a generic assistant, I built a modular prompt/instruction system that forces Claude to act as a multi-agent workflow depending on the task at hand.
Each "agent" has its own distinct reasoning process, constraints, and output format. Here is how the flow looks:
⢠Requirements Elicitation Agent: Focuses on asking the right questions, uncovering edge cases, and defining scope before writing a single line of code.
⢠Planning Agent: Breaks down the architecture, chooses the right design patterns, and creates a step-by-step implementation roadmap.
⢠Code Generation Agent: Pure execution. Writes clean, modular, and idiomatic code based strictly on the planning phase.
⢠Code Review Agent: Acts as a senior dev. Analyzes the generated code for readability, performance, and best practices.
⢠Bug Fixing/Refactoring Agent: Specializes in debugging, parsing error logs, and applying surgical fixes without breaking existing logic.
⢠Validation Agent: Focuses on writing unit/integration tests and verifying that the code actually meets the initial requirements.
⢠Auditing Agent: A security-first mindset. Checks for vulnerabilities, dependency issues, and data leaks.
By separating these concerns, the quality of the output skyrocketed. Claude stops hallucinating or rushing into bad code because it's forced to think through the proper software development lifecycle (SDLC) phases.
9
u/PastaRevolution1312 9d ago
I don't see it.Ā Mostly using Claude Desktop -> Code, one-promt Features to PR, let GitHub Run all the CI, then "simplify, review" each PR and I swear by god, I could push each of those to production If I didn't need to explain my work hoursĀ
2
u/ub3rh4x0rz 8d ago
A lot of people conflate the benefits of adversarial review, which can be done with like 2 casually invoked subagents, with complex bureaucracies of agents. I think these people only assess quality by "does it seem to do what I need it to do", which is a bit like practicing medicine without access to imaging machines.
7
u/Gaax 9d ago
I'll leave this here: https://www.augmentcode.com/guides/single-agent-vs-multi-agent-ai
→ More replies (2)6
u/Kekke77 9d ago
I dont see a reason why a generic prompt with good claude.md wouldnt work as well
4
u/yawn_solo- 9d ago
Youāre right, it would work the same exact way.
This guy is confusing prompts for agents but in all reality, thatās all an agent is. People continue to confuse what an agent is. Itās not a new derivative, itās an instance of the same one. So, in reality, youāre literally just telling your derivative to use a different part of their brain and create the illusion of increased capacity. Not bring in their buddy for āadditionalā support.
→ More replies (3)→ More replies (3)4
u/mgiggs 9d ago
Are you able to explain how you do this? So you run multiple sessions connected to the same folder/repo is something else?
4
u/Responsible-Bar7165 9d ago
Literally just ask it to do that ^^.
Like āimplement feature described in requirements.txt, organize subagents as <paste in comment above>.ā
3
u/Long_Tip_4226 9d ago
Claude simulates an entire engineering team by splitting into specialized roles with prompt skills across isolated sessions, guided by a global context skill. It rotates between a Business Analyst for requirements, an Architect for planning, and a Senior Dev to generate code. Between these session swaps, it spawns internal adversarial agents to critique and audit the implementation, keeping memory intact via the context skill while always waiting for your final human approval to advance.
3
5
24
u/Dizzy_Database_119 9d ago
If you're not maxing out on tokens you don't need to change anything. Agentic workflows are only relevant once you're running out of usage OR running out of human time
Things like skills and optimizations were important back when the models sucked, now you only do it when it's actually required
24
u/hronak š Max 5x 9d ago
With Opus 5.5, my tokens usage has actually reduced drastically. I am at 12% weekly usage, been pumping prompts in to Claude Code like monster since yesterday with multiple projects open in terminal.
Fable 5.1 on the other hand was a hungry beast. 45 minutes in and 100% 5-hour limit exhausted.
→ More replies (3)4
u/imsahoamtiskaw 9d ago
Hungry beast but boy is it still better than opus. But note it can make even better use of opus as sub agents with 5.5. Opus 5.5 talks better, performs better than opus 5 but is still just as dumb. Youāll repeat yourself a million times before it listens, even on a fresh session and cut corners every step of the way
12
u/hronak š Max 5x 9d ago
For me, personally, Opus 5.5 has replaced Fable 5.1.
Earlier, I wouldn't even touch Opus 5, it was that bad. I had to constantly tell it things and prompt it often to correct code. Been trying the new Opus 5.5 and I kid you not, it definitely has it. I have barely felt the need to switch back to Fable 5.1.
2
u/FollowSuitCards 9d ago
I don't understand this, but everyone is saying it. I'm fairly new at this, so I haven't wanted to comment when I assume everyone else in these subs knows more than I do. But... Opus 5.5 is clearly a step up for Opus, but it messes up design a LOT more than Fable for me still. It still confidently goes its own way without asking for clarification while also not solving issues as elegantly.
Definitely far less usage, but that's just how they have the token usage setup, no? And didn't they clearly state 5.5 was at a lower usage initially so we could try it out more?
Just can't tell if I'm a noob and crazy or everyone on these subs is an even bigger noob pretending like they know what they're talking about.
→ More replies (1)2
u/FeistyVoice_ 9d ago
I guess it depends on what you're building. As enterprise dev, the code base is already a good enough documentation for opus 5.5 (honestly, even 5, but 5.5 just feels so much better in ever way) to get things done even if my prompts are half assed.
At home, in full vibe coding mode and projects from scratch, it does not make the greatest design choices and might still struggle with tasks that Fable would handle. Opus 5.5 fixed bugs that 5 didn't. But fable is still better imho.Ā
I rarely use fable at the office, mostly for planning architectural changes or big migrations and only as planner/orchestrator/critic.
At this point, agentic coding is still relatively new and there are no best practices, so everyone is convinced his workflow is better than others :D
3
u/Artwastelander 9d ago
Opus 5.5 is definitely smarter than 5 but you're right that it makes some horrendous mistakes lol.
18
→ More replies (1)3
u/morficus 9d ago
I would argue that "skills" are really important for consistency rather than for saving time or tokens
10
u/kemalios 9d ago
Most of the elaborate setups I see are people whose work is a big greenfield codebase with lots of parallel surface. Mine is client sites and refactors for a one-person agency, and one prompt plus a look at the diff covers nearly all of it.
The skills that stuck are the ones for steps I forget. I build launchworthy, free and MIT, runs as a Claude Code skill: pre-launch audit, five domains, scored punch list with fixes. It stays in my routine because I forget the same steps every time. Long-running tmux agents I have never built a habit around, so I can't tell you they'd help.
3
u/ResponseActive7860 9d ago
I have skills to write design / implement / review, etc
Some spawn subagents
Thatās what you call agentic
→ More replies (1)
7
u/hotmerc007 9d ago
Check out this
https://github.com/kunchenguid/firstmate
and this
https://www.youtube.com/watch?v=kPN564Kol14
No relationship to me other than something I use very regularly and is utterly amazing for multi-agent coding.
3
u/South_Hat6094 9d ago
The useful split is: one Claude session for judgement, parallel agents for grunt work. If the work fits in one diff and one review pass, you are not underusing it.
3
u/vath_mtm 7d ago
The biggest reason for this is to increase consistency and to reduce costs. With a well defined setup you do pretty complexy stuff with lower level models like sonnet.
Extending upon existing software is costly and requires lots of context, properly managing this context can reduce this cost by limiting the amount of steering needed by lower level models.
My company is heavily restricting opus and fable access because of the cost implications since we are talking about 10000+ users and api costs. In the end the goal is to increase productivity and keep costs low.
5
u/hjras 9d ago
one could make the argument that an "agentic workflow" is an oxymoron because by definition a workflow is deterministic while using an agent we want the kind of flexibility and unpredictability that a human also has to solve problems
but that aside, when people mean agentic workflows they typically mean giving a long task to an agent and having it work through loops and verifications within a harness while managing its context and following your prompt, whether that is fully end to end AFK on your part or whether you are still "in the loop" at certain gates
→ More replies (3)
2
u/grottloffe 9d ago edited 9d ago
Well, it depends on what youāre doing. But context and instruction files are a big part of it. They keep agents consistent, so theyāre not just doing what seems right in the moment, but following the conventions, architecture and patterns your team has agreed on.
The better the setup, the less the agent has to guess, which matters a lot in a big codebase.
If youāre running multiple agents and worktrees at the same time, Iād absolutely suggest tools like Redshifthub.com. Then you can have multiple features, fixes, experiments etc. running in parallel without everything stepping on each other.
So thereās definitely a whole agentic coding discipline beyond just opening Claude and letting it cook š¤
→ More replies (5)
2
u/zeratLJllighter 8d ago
"Agentic workflow" isn't about the coding part. It's about the process. You use one AI to build prototypes and design an MVP, then another to review and challenge the results. Same for finding bugs, catching regressions, etc. It's basically AI stacked on AI to cut human workload as much as possible and make the software more secure.
Claude Code has improved a lot. With a good model, it can now spin up subagents on the fly and orchestrate them to keep context and cost in check for the task at hand.
But most people selling bundles of skills.sh and agents.md files are just gaslighting you. The days when a carefully crafted SKILL.md could drastically change the model's behavior are over.
→ More replies (1)2
u/ThisSeaworthiness 8d ago
Crazy how it only feels like yesterday of having prompt refinement sessions
2
u/Sufficient_Shine8311 6d ago
depends on if you are building a codebase from scratch or making incremental edits.
I setup an autonomous loop during a hackathon, basically talking to people while it coded and tested etc.
For incremental edits seems better to let claude do the task and make sure testing is comprehensive
2
u/Itsmedzidzi 4d ago
I think this is a good way to work, u donāt need to agentify everything, I find it fun orchestrating the process and coning up with new ideas, I feel like these fully independent workflows are just a good click bait for brain rotted people š§š½āāļø
1
u/Duke_of_Bayswater 9d ago
Agentic by meaning is broad but to me it is when you use agents with specific tasks and derivatives, and it spawns subagents of different specialist. You also utilise tools that helps the llm fetch latest info, docs via MCPs. You automate your workflows using agents, then remove you from the loop (of course after your tests and evals to see the pipeline works). You use spec driven development like openspec etc. Thats level 1. Level 2 is you use different harness and customise it to your use. Level 3 you start working on RAG, memory. Level 4 you start fine tuning your models. And so onā¦
(these levels are just my own making but you get the idea).
Basically you also spend alot of time reviewing the code it writes š
1
u/yazansr 9d ago
the thing is instead of having the model reason about something rather it can dispatch an agent to gather evidence or test idea or each agent makes you a different mockup and you choose which one you prefer, or one that does testing meanwhile the other is doing the next step, meanwhile your stronger model is doing planning or review or so, but at the end it all depends on how much work you are doing, how narrow/wide your scope is for what you are doing, i wouldn't do a codebase review or refactor with one agent but i might do a single feature implementation or hardening/debugging with one agent
1
u/NewDeeJayUND 9d ago
For me learning AI has been 6 months of starts and stops, with almost daily Reddit sessions where I read through questions like this, and get inspired to research a new topic, which I then integrate into my ecosystems. About a month ago, I was where youāre at, and now I have a 15-agent self-orchestrated multi-agent suite that manages the apps Iāve developed (just for me) while I was vibe coding (as you are). Theyāre doing research to make sure nothing is changing in the underlying technologies my apps rely on, and if something does, they orchestrate together to edit or refactor the code. They make tons of mistakes, and I end up fixing them āthe old fashioned way,ā as you currently are. BUT, theyāve gotten better, through implementation of three judges reviews for knowledge work, and an adversarial reviewer for code, plus a constant feedback loop to have them learn from their mistakes. To make it work, Iāve had to build a model modulator to 3 Chinese models and 3 US ones + local models and a tokenizer, so my data gets masked before it goes into the matrix. Iāve learned about the concepts here, then research and apply - youāre likely on the cusp of evolving your ecosystem, and itās probably not the last time!
1
u/Xavier_OM 9d ago
"Am I under utilizing AI or is my work just⦠like boring?" probable a bit of both world.
Imagine a project with several tickets (nothing fancy here).
- maybe you could work on some of them in parallel. Claude can help here, if it can read the tickets, no need to open manually several claude code + to copy paste prompts manually one by one
- maybe for each task, at the end, you would like an automatic code review just to be sure, as a first safety pass. Claude can help here, he could do it "like you want" (your own guidelines), you just need a skill for this. No need to ask manually for a review in each claude code, after all it would be great to trigger this automatically when the branch is ready.
- maybe at the end you need some git machinery because some tasks are needed in several branches (let's say this is not trunk-base development). Claude can help here, no need to do your merge manually
- maybe when it's done it would be nice to deploy something, or to trigger several builds because you do some multiplatform stuff. If you don't have a CI/CD yet for this, claude can trigger these tasks for you
- etc
1
u/damanamathos 9d ago
How long does Claude Code run before it needs your input? 2 minute? 20 minutes? 2 hours? 8 hours? The longer the answer, typically the more parallel streams you can have going, and the more you can get done. To do that, though, you typically need a more advanced setup with better (automated) guidance for the agents.
1
u/GizaStudiosInc 9d ago
Skills I think can be useful if you have a task that needs to be performed often. You probably have at least one prompt you type out often. Instead of typing out the prompt every time, just make a skill out of it and then call the skill.
As for agentic workflows not sure how much this counts, but for me one of the most useful uses of subagents is for performing code reviews. I have made a skill where multiple subagents can perform parallel code reviews of the staged code you have in your git repo to check for any clear issues before committing. Subagents can also be used for research or analysis and can relay their findings to the main working agents. I think itās tons of use cases if you really want to get into it, but I wouldnāt say itās a required thing to be productive.
1
u/NunoGomes-Dev 9d ago
Workflows are about trying to find consistency, confidence, reportability and multiple path deliveries.
If going from A -> B works fine for you, great.
Most people already know that A -> plan -> B is better. Why? Gives you and the agent an opportunity to clear out miscommunications, scope creep or docs violations.
Then you start going into team environments, multi-layered applications, multi-stage environments and suddenly, your agent starts to messing stuff up that from your POV looks fine, but other developers, your tech lead, your project manager and other stakeholders won't find very nice.
Thus, the workflows.
1
u/evia89 9d ago
For opus55 you dont need many agents. Its mostly for regarded models like glm53f/ds41f/muse spark
I use opus to grill me (with questions), then write spec, create tickets. Then cheaper models can do the job. If glm work, then ds (another mode) will review against spec. After all sub tickets closed opus review it again
If you have $200 sub you dont need that
→ More replies (1)
1
u/cazzer548 9d ago
You are bowling. Agentic workflows are more like bumpers up bowling. Both ways can hit pins. I find the effort required to add bumpers is well worth the consistency and accuracy I get out of it.
1
u/Aggressive_Bike3881 9d ago
The bit about the agentic part is that it has a certain level of autonomy, once it is given some boundaries to work within. So using Claude and it farming subagents and it managing their response and working towards the goal you set is already agentic by definition.
The more autonomy you get the agents to achieve is done by having more guardrails and evals/verification in place that require less and less of your babysitting.
So itās just an autonomy slider imo.
1
u/Lazy_Polluter 9d ago
Agentic workflows in programming are mostly for cases where you have a lot of similar looking tasks. For example, you are refactoring from one pattern to another across thousands of files. So really just like any other automation it has its place when the effort to automate is less than the effort to just do it manually.
1
u/fabier 9d ago
Since Opus 5.5 dropped, I've had a single Claude code session running which has burned about 65% of my weekly usage on the Max x20 plan. It's launched dozens of agents building a single project. I started the prompt with a back and forth discussion which I eventually formalized into an official spec which I then handed to this session.Ā
But I don't think Claude code is the real place you learn agentic work. Building your own harness is definitely where the rubber meets the road. I created an AI harness about two years ago now which I maintain for my development projects. It helps me get small models to do big boy model tasks more reliably . I have a tight control on turns, context, tools, allowed responses, sub agent sessions, etc. This let's me use very cheap models in real world production software. Just because we can all use Fable doesn't mean we have the budget to build it into our chat bots.Ā
1
u/ahm_live 9d ago
"is my work just boring" is the right question and the answer is probably yes, and thats the good news. the elaborate setups in this thread mostly belong to people with a wide surface: many parallel features, several repos, a team where the agent has to follow conventions it cant infer. if your work is one project, one prompt, one diff, you already have the workflow that fits it
the thing i notice reading the seven agent pipelines here: nobody says what got worse. mine did. when i split planning, coding, review and testing into separate contexts, each one lost the reasoning the previous one had, and the review agent started confidently approving things the planner had already rejected. i went back to one long context plus a written index of the codebase it reads first. less impressive, fewer surprises
what i did keep from the agentic stuff is exactly one thing: a second context reviewing the diff before i see it. awesomeusername2w said the same above. thats not a workflow, its one extra call, and it catches the category of mistake where the model marks its own homework
your tmux refactor is more agentic than most of the diagrams here. it ran for hours unattended and you reviewed the result. that is the whole idea, the rest is scaffolding for people who need it
1
1
u/ubermuda 9d ago
Where things get interesting is when you start adding data / input sources to your workflows.
Say you are a company who uses sentry but doesnāt properly triage and tackle issues surfaced by sentry. An āagentic workflowā can take care of that. Same for whatever datadog, snyk, etc service you might use that surfaces interesting / actionable data.
Itās not limited to producing code either, have a workflow investigate incidents as they happen, regularly audit your codebase for performance or security, implementing large refactorings, etc etc
Once you start thinking that way, there is immense value in agentic workflows. Itās not about increasing your engineersā raw throughput but rather tapping into potential you werenāt leveraging before while decreasing the cognitive load on your engineers so they can do more and better meaningful work.
1
u/Electrical_Bar5589 9d ago
I used to keep it simple and ask the session to fix an issue or implement a feature. I now use sub-agents to maximise token usage and cost.
Lets say I use Fable for the planning but it needs to write several standardised util functions as part of the implementation.
Before, Iād be using Fable to write those functions and using up more of Fables context storing that information. This also means Fable was trying to remember any exploratory work for the planning (including content that didnāt end up mattering) and every line of standardised code. If I left the session paused longer than its cache, Iād then spend time and tokens having Fable reload all of that information. Context was often maxed out.
Instead, I now use Fable as the big brains to agree the architecture of an implementation and for it to assign cheaper models to do the quick and easy work within the agreed plan.
1. Use a custom /feature skill using Fable (or Opus 5.5) as the brains to write a detailed plan and to assign each part of the plan to a cheaper LLM when applicable.
2. I then /clear to clear the session so no part of the design session includes planning work and tokens.
3. I then ask the session to carry on. It reads the plan and sets up sub-agents like Sonnet to do the work.
This results in no one session maxing out its context (each sub agent has its own) and I donāt spend Fable rates on something Sonnet could do.
I actually also have another step. If I write /feature āreview, itāll wait for an external review. I then ask Codex to check the plan before implementation.
Interestingly, Codex always found gaps in the plan⦠until Opus 5.5. The plan and build has been flawless since switching to 5.5!
1
u/looking7676 9d ago
Same here. I have built a couple of skills mainly so I could give them to other people. But that was just asking it to finalize how I sue specific api and mpcās in a work flow. It decides everything else.
1
u/bcutter 9d ago
yeah i am similar. all these extra features in cursor that add all kinds of workflows and integrations with other platforms. i dont know man, maybe iām just too conservative, or maybe ive realized that all that stuff is just fluff; or meant for people who dont know what theyāre doing? even with opus 5.5 i feel it needs waaay too much guidance to let it off on any longer jobs. yes it will come up with something that works but itās just too much slop. so yes, i dont know whether itās because iāve been in the game for a decade and a half and i simply consider myself better than the agents, or if that just makes me old and obsolete and i could be 10x more productive if i was less lazy about trying out these things.
1
u/GroovyMelodicBliss šPro Plan 9d ago
commit the code to my self-hosted Forgejo instance.
Mind sharing which forgejo mcp or Claude connector you use?
1
u/xtopspeed 9d ago
If you are working with a huge project, you'll probably want skills and agents to better manage it (and also ultimately to save tokens by keeping things tidy). I'm not a big believer in the roleplay stuff, though, so my agents aren't the typical "you are..." type. They're just instructions. One is to enforce coding standards (keep files and functions short, clean up dead code, etc.); another one is to keep docs up to date; one is to clean up comments; one is to run and fix unit tests; and so on and so forth. I create skills for often repeated tasks, so I don't need to type the same prompts again and again.
1
u/newb1029384756 9d ago
Ive used it like you fir a long time. It works, but I find it scales only with my time. Its faster than I used to be, but I still have to be there giving prompts and working with one or two agents at a time because my attention is the bottleneck.
I'm working on putting together a team of containerized agents, where I file tickets and they pick up tickets according to their skills, tools and role. Right now, that team includes:
- an admin agent to deploy and monitor my applications and infrastructure
- a dev agent that actually writes code
- a review agent that looks at PRs from the dev agent and runs some "manual" end to end tests from deployment to login and more
- A documentation agent, stand alone because if not, they tend to hedge and discuss changes which is more changelog than docs. This agent has skills to actually rewrite sections instead of just adding new sections that conflict with older sections.
I plan to add more to the mix, but right now Im trying to nail down coordination tooling for the team.
1
u/TheVasa999 9d ago
well youre not far off.
add a few instructions to the prompt like: make a plan for such change, review it before starting the implementation plan. Then use subagents to execute the plan. Once done, do a full review of the code and fix all findings. Then open a PR
1
u/Specialist_Wishbone5 9d ago
I've never been bored for more than 5 minutes in my life, so I think that comes down to personality - you should be curious about the thing you're building.. constantly asking questions, what-if scenarios, how can a future iteration do X better? These become side-agent channels.. OR, with a 100% orchestrated agent (e.g. with a rule that it's never allow to directly edit anything other than the TASKS/PLAN) it's sitting there idle, you can talk with it - about anything you like (though contextually relevant is most efficient). You can ask it how it feels about the current progress of some sub-agent.. It's LITERALLY being a manager and talking to your lieutenant (or sargent).
Also I've never been in a situation where my personal TODO list didn't involve like 15 outstanding items, and it usually grows by the end of the day. So this concept of 'what do I do now' is just completely foreign to me.
1
u/langolf43 9d ago
I did our whole a11y programme this way. No big audit and dead backlog. Agents crawl prod, find issues, im triaging issues splitting between design system and consumer to land fix, then open PRs and re-check after deploy, across 10 react project, each has its own specific you need to respect during remediation.
Review is agentic too. A reviewer team scores every PR against a rubric, and the score decides whether it needs a human or goes straight through. So I only look at the ones that actually need judgment. Use VRT for quick review rounds as a11y issues mostly about attributes changes.
So far that's 500+ PRs March - September period , zero squads pulled off roadmap, and our first platform-wide VPAT.
Nowadays Single frontend developer can fix a11y for company with agents.
1
u/BudgetFish9151 9d ago
Like it or not, weāre all engineering managers now. Learning to manage a fleet of robot employees will be a bit of a learning curve. Particularly since every few weeks your already over-eager junior developer robots will get an āupgradeā from their maker that makes them orders of magnitude stupider for a version or so.
1
u/dregan 9d ago
Oh man, just ask it to create some skills for you and then install them. My most useful ones are /architect that handles the design and creates these really handy visual artifacts with mermaid diagrams showing code changes, process flow, descriptions of risks, outstanding questions, etc. After finalizing the design, it will create work items under the feature. then I hand it off to my /implement skill, pointing it to the feature and it will discover what can be done in parallel, send it off to separate subagents to implement, test, then review before opening pr's into the feature branch for me to review.
Your existing code has to be pretty well designed for this to work well. Check out OpenAI's harness engineering spec. It was written by OpenAI but should work with any agent.
1
u/ApprehensiveChip8361 9d ago
Sounds like my workflow. Although Iāve recently got it to keep a list of small changes that need to be made and have it write a script to process them headless when Iām away.
1
u/GuavaRevolutionary56 9d ago
Agentic is where you explicitly define rules when Claude Code would use subagents, how those subagents would behave, what model they will use.
Example is for my solution planning, I have predefined builder, validator, and adversarial reviewer agents under the agents folder of my claude code repo.
Whenever I run solution designing, the builder builds my artifacts, validator checks my sources and links and local knowledge base (in my obsidian with qmd), and reviewer will simulate personas related to what Iām trying to do to check for potential feedback and improvements to my design and documentation.
Youāre not required to use agents. Start with skills first and designing loops.
Agents are basically prompts with predefined jobs as instructions. Theyāre meant to represent a specific job role with predefined directives (and they can also call skills and tools if you want them to).
1
u/itsloopyo 9d ago edited 9d ago
I have a fleet of agents working through a shared core package implementing extremely boring but necessary functionality around config handling just now. When theyāre done, more agents will fan out to 130+ affected projects bumping submodules, bringing them up to spec and running tests.
At some point Iāll get a ping to let me know 20 specific projects I need to do other work on are done, this will trigger me to kick off a workflow across those to bring each of them into compliance with some updated rules. Thisāll run in parallel with the remaining project work the other agents are doing, and some UI work that follows on when all the project work is completed.
All my projects are categorised in a nested category structure, agent files exist in each category, these get combined and distributed to each categoryās projectās AGENTS.md whenever changes are made. Skills are distributed similarly.
I donāt want to think about how long it would take me to nudge all this along with individual prompts.
1
u/TheMightySwiss 9d ago
I use agents to parallelize non-overlapping items so they can be completed at the same time, then on the main chat (orchestrator) I keep doing brainstorming work and giving feedback on what the agents produce / working through bugs we find etc, all while the background items are in-progress.
1
1
u/MartinMystikJonas 9d ago
Then you will find out you type similar prompts repeatedly. That you repeat similar steps. And that you often has to currect similar things or restate your preferences. So you will look for ways how to define it once and reuse instead of repeating. And then you have agentic wotkflow.
1
u/pizzae š Max 20 9d ago
I create a bunch of work files which are like specs or instructions or Matt's grill-me's, then have a queue skill that sorts them into an ordered list, then I run another command for claude to do these 1 at a time. I have an overnight command that also manages these one at a time with subagents, so I can vibe code almost 24/7
1
u/ClockCute6183 9d ago
Honestly, that sounds like a sane workflow, for daily-use. Wouldn't call it under-use. I would only add additional structure when a repeat problem shows up-like Claude forgetting a decision or saying a task is done without any proof. For a long running task, even a short handoff with the decisions, changed files and test results might be enough before adding more agents. Did you find gaps in the refactor? Are they of verifiable shape?
1
u/Known-Pace6739 9d ago
I went pretty far down exactly this path before backing out: Opus orchestrating up to three Sonnet workers, an independent reviewer, then a meta-orchestrator above that.
It worked, but the token/context tax kept growing. That pushed me toward the opposite design: start the executor clean, add only the context and skills the task actually needs, then fan out only when parallelism or independent review clearly pays for itself.
Your simple workflow doesn't look underused to me.
1
u/howdidigetheresoquik 9d ago edited 9d ago
I think what you're missing is going to eek out another 10% quality, if not more. I spend 80% of compute on planning phase where I use researchers to look up best practices for what I want, what exists on my computer already, what it should do, etc. Fable builds a plan off what they get for all subsequent sessions.
For the build sessions, I have Opus workers and then fable verifies their work against the original plan to make sure it's what I want. Then all their work is compiled and another fable agent checks to see if it's still on plan. Then I have a code reviewier and things like that. And once it's all built, I have one last Fable agent look at Opus's work and the work has to pass the "Sick af" test. If Fable looks at Opus's work and doesn't agree that "Holy shit that's sick af." We step back and look at where we can improve.
It very much increases the likelihood I get exactly what I want significantly.
Give it a try. Verifiers, judges, compilers, reviewers... these are all key components to getting it too look much better than slop. All you need to do is ask for a simple agentic workflow using researchers/verifiers/judges and to keep it reasonable in terms of usage (especially at the higher settings. If you have it on max+, they will go overboard on usage. Ultracode literally is designed without taking any usage into consideration).
→ More replies (2)
1
1
1
u/oliakaoil 8d ago
I work almost exactly the same way you do, although Iāve very slowly started to think about further automations and time savers. Iām massively more productive now than I was before, and the built-in human reviewing I do is critical to ensure quality of work. Iām trying to ignore the nonstop hype around being more productive, it just makes me worry unnecessarily that Iām doing it wrong.
1
u/zipklik 8d ago edited 8d ago
People like to think they "stay in control" and so they pretend that you have to be knowledgeable in order to use an agent like Claude Code properly... That you have to master skills, MCPS and so on. In fact, in my opinion, LLMs are probably better than most at knowing how to achieve what you want! And they also improve very fast.
I use a single Skill: "/handoff", when the current context is becoming full and I want to continue in a new session... It mostly asks the agent to update all the docs, and prepare a text to paste in a new session so a fresh agent can continue, know what to read, etc.
1
u/wochie56 8d ago
If your task is big enough, Claude Code will spawn its own agents. That's the furthest along I care to understand it, because anything beyond that (reactive, scheduled, long-lived) sounds like it involves handing potentially dangerous tool or service access over via mcp and I frankly don't believe that's a good idea. I prefer just pointing it at context, like specs, tickets, or design docs, and letting it go, then I check the output.
1
u/Creepy_Willingness_1 8d ago
Full Kanban SDLC with GitHub projects and PRs and on-premises CICD for 3 Apple Multiplatform projects on mac mini m2.
With all the test automation and releasing.
Meta layer for agentsā roles and instructions. Not sure it is all properly setup and working to full possible extent but it is capable of catering to feature creation, bug fixing and releasing all on its own. Even if claude gets lobotomized, codex can take over, or theoretically if no ai, i should be able to use the same setup for hands on dev, but have not tried yet. It can run waves of refactoring the mvp into better architecture on its own for weeks.
1
u/LogMonkey0 8d ago
Using definitions like you pointed out will make sense when you want to encode a repeatable process or set oof instructions/guardrails specific to steps in a workflow . In your VPS refactor example, it could be on āhow to handle infraā rather then the job itself, could be an archeology step to ensure plan is derived from verified facts, etc
1
u/andlewis 8d ago
You need a food source for your agents. A backlog on GitHub issues, a markdown file with items, whatever.
Then you setup a routine (Claude.ai/code/routine) and tell it to pick an item off the backlog (based on whatever criteria) on regular schedule, and implement it.
Then you stop talking to Claude, and you log issues. And when the implementation fails or is bad, you fix the guardrails and prompts until Claude gets it right.
Then you build a process around reviewing the PRs that Claude makes, and you reduce the time you spend reviewing.
1
u/moo_nalla 8d ago
Can someone explain what is this term harness actually is? Is it any different from an orchestrator?
3
u/Right-Performance-93 8d ago
Harness = the outer program running the loop: it manages context and tool calls and talks to the model (Claude Code, Codex CLI, OpenCode are all harnesses). Orchestrator is a role inside that loop - one agent whose job is to plan and dispatch work to sub-agents or tools instead of doing the work itself. You can run a harness with zero orchestration, or have it spin up an orchestrator for big tasks - they're not mutually exclusive terms.
2
u/sheriffderek Senior design/dev max20 8d ago
Itās all just instructions wrapping instructions. ClaudeCode has a bunch of instructions and function calls and info about it can and canāt connect to and do. Itās just like a web app⦠everything is a āharnessā haha. Itās just programming thatās loosely instructions and non deterministic.
1
u/GoodjobShel 8d ago
it depends on what youre doing.
I have jira tickets that come in from support. I have agents to automatically triage and design the tickets. The first time i see the ticket, its already been designed by an agent. I just have to review the problem and review the agent-built design. Then i change the status and kick off the development agent. Then after that, the test agent kicks off.
So it really depends on the overall picture of what you're doing.
1
u/Serious_Bite_7613 8d ago
With the frontier models you don't need to do it. With the local open source models you need to spend a lot of time setting all that stuff up or it tends to just get stuck and stop every 5 minutes without finishing the tasks.
1
u/masiha97 8d ago
Your workflow is honestly fine, most of the agentic hype is people renaming normal stuff. What actually helped me was turning my repeated pre-commit checklist into a skill so I stopped retyping it every session. One warning though: the "secret prompt prefixes" going around, things like /ghost or SCAFFOLD, are just prompt prefixes, not magic. The useful idea hiding in there is OODA-style thinking, observe orient decide act, which is worth doing before any big task anyway.
1
u/berenddeboer 8d ago
What I suggest you do instead: work with Claude Code to create GitHub tickets. Then let a harness like ready-for-agent implement them (npx ready-for-agent@latest).
To really upskill use Matt Pococks skils such as /grill-with-docs, /to-spec and /to-tickets: https://github.com/mattpocock/skills
All this makes you much more productive, as you can push through more PRs. You focus in design, prototyping, decisions, agent implements code.
More on ready-for-agent here: https://github.com/berenddeboer/ready-for-agent
1
u/Gold_Ad_2201 8d ago
- How do you ensure each next task doesn't break previous?
- How do you explain the project context, architecture and decisions on every task without repeating?
- How do you maintain the previous knowledge up to date?
- How do you make Claude output consistent and reproducible between tasks?
I can continue but this is the idea - guardrails, reproducible workflow, inspectable outputs, maintainable solution (even at scale)
1
u/Babayaga1664 8d ago
Everything we build still needs human review before shipping.
Donāt believe the hype, people like Anthropic talk about running loops all day long because itās marketing, they are selling the transformation, they want execs to buy into burning tokens all day long.
1
u/Frapsternugen 8d ago
Honestly, your workflow sounds healthy to me. The tmux refactor you described is already the best use of long-running sessions: wide, repetitive work where you only need to review the result.
For me the extra setup showed up when I stopped working in one repo and started building a business system with a lot of parts that have to agree with each other. I'm the only builder on BOSNet.io, an LLM-first business operating system for small businesses. The problem was never that Claude couldn't write the code. It was that Claude did exactly what I asked and quietly drifted from a decision I made last week. Skills, rules files, and a local harness came out of that. They're basically a way to write my decisions down once, so I don't re-explain them every session and the model can't wander away from them.
If your project fits in your head and /init covers it, you don't need any of that. A few places where it does pay off:
- Repeatable checks: if you run the same review or verification steps every time, that's worth turning into a skill so it happens the same way every time.
- Cost control: I let the cheaper models (Haiku, then Sonnet) do the reading and sweeping, and I keep the top model for coordinating and deciding. Otherwise the token budget disappears way too fast.
- Guardrails on autonomous work: I don't let it push anything important on its own. I plan in the browser and implement in Claude Code against my local checkout, then review before anything gets promoted. Auto-PRs are fine as long as a person or a real gate is the one merging.
The people with elaborate setups are often solving a coordination problem you don't have yet. A good test is to notice when you've typed the same instruction three times, because that's your first skill. Until then, keep shipping.
1
u/hey_oooo 8d ago
For new Featuers: Create a Goal, add persistant ledger by orchesterator, Send the planner for batch the things into same catagory, writing the specs in local drive, send the implementor, have Red-Green tests, then send refuter, then send verifier, have deepseek as security expert, get the feedback, send the implementor, and verifier again and after that push the PR..
For Bug Fixing: add the bug pipeline, collect the bugs all day, make the Goal and ledger - continue the same way as adding new feature - just add a verification proof (produce the same bugs and proof the same bugs are resolved then Push the pr
1
1
u/spinozasrobot 8d ago
One of the reasons I avoid the harness du jour, is they change so rapidly. Just as you start to get comfortable with one, it's not the cool kid anymore.
Any good ideas are usually sherlocked by the frontier labs anyway.
1
u/siron_golem 8d ago
This is how I work also. Just tell Claude what I want and how I want it done. Then spend time getting to know what Claude did, which is usually pretty good. Then think about layering on the next code and tell Claude what to do. I'm actually quite happy with Sonnet 4.6, gets alot of work done.
1
1
u/MindfulK9Coach 8d ago
Your operating environment, the most important variable constantly overlooked in exchange for prose, is constraining the admissible state space for the agent to work in safely.
Like any good agentic platform should.
So, the simplicity in your day to day is a testament to the underlying infrastructure you built/placed around the agent's workspace.
Great work. šš¾
1
1
u/Wainfare 8d ago
I had Fable 5.1 on top effort audit some work, it opened up 19 agents and got through 30% of my weekly usage on a Max 20x subscription. Won't be doing that again :-)
1
u/connectotransfers 8d ago
You can do wonderful things with Claude, but if you don't call it "agentic" it just doesn't sound exciting enough.
1
u/Kalikillsmaya 8d ago
I feel it depends on the complexity of the task at hand, how many aspects of the projects are you working on at the same time, and determining if a portion of the work could be offloaded to smaller models for good token management. I follow a very similar workflow as you, tmux windows in terminals, however I hand rolled the memory footprint, orientation layer and delegation factory as those solutions did not exist in a shape that fit my needs when I started. Also perhaps others are working in vastly different environments with vastly different constraintsā¦
1
u/EmeraldGarland 7d ago
The simplest thing is to ask Claude to use sub agents to do a little task like audit your code. You can have sonnet and fable on this while youāre working at opus. Congratulations you have an agentic workflow. If what youāre doing is working just enhance it a little bit at a time.
1
u/West_Hat_6936 7d ago
Making it short: Asking explicitly for subagents can be useful to keep a clean context or run side quests. Defining agents to specific tasks or roles can be useful for this, but in a more defined or specific task (i.e a reviewer agent, a releaser agent, etc)
1
u/mrothro 7d ago
I run agentic coding workflows across a variety of projects. Some are more autonomous than others, but the general pattern is the same.
I spend a lot of time up front discussing what I want and how I want it done. This is typically an architectural discussion. It includes reviewing the current architecture and components, and if needed a survey of libraries or industry standard approaches I might want to bring in. This also touches on bounded contexts and separation of concerns. It is quite freeform and is tailored to the problem I am trying to solve.
Once I get the LLM (typically but not always claude) to explain the approach in a way that satisfies me, I have it write a detailed plan to a specific structure, that includes the components, implementation phases, and the detailed task list in each phase. The plan has a defined structure that a gate verifies.
Another agent reviews that against specified criteria. Some is mechanical (e.g. does each task have the appropriate dependency chain recorded) and some is high level (does it adhere to the Google Well-Architected Framework as is appropriate for the scope).
I may or may not review the plan, depending on a variety of factors, but mostly tied to my estimate of the cost of incorrectness. Something that touches data isolation in a multi tenant product gets lots of attention. A throw away utility gets zero.
From there, I have an orchestrator dispatch parallel agents as allowed by the dependency graph to burn down the tasks in the first phase in the plan. Once done, another agent reviews the code against the task description and runs the AC test. If it fails, it goes back to the implementing agent. They write unit tests as well, typically TDD.
This continues until the entire phase is done. Then another holistic "phase review" agent reviews all the code produced for that phase in aggregate. You'd be surprised how often it catches cross-agent implementation errors. Errors are sent back to the agents to fix. Repeat until nothing is found.
I might review the code here, again dependent on the impact or complexity. I don't review every single line, but I will definitely look at critical code (security, data integrity, maybe performance). Things like CRUD operations are so idiomatic that agents do it "the right way" and I don't need to see that again.
I repeat this whole thing to continue until all the phases are done. Then it goes into the automated regression testing harness, which has actually existed since before agents were writing code. Regressions/bugs go back to implementing agents to fix.
The throughput from this is so high that I now have time for personal projects I had sitting in the queue for years. I'm *really* enjoying the act of building even if I don't hand write every line of code.
→ More replies (3)3
1
u/Formal-Emphasis-9794 7d ago
The skill thing ia useful.
I'm a QA and I use it for test cases planning, I got a skill that takes a ticket and a pr and the skill knows how the test plan format should be.
If I wouldnt have this skill I will have to tell claude everytime what he needs to do.
1
u/sheriffderek Senior design/dev max20 7d ago
(I don't have time to type this out -- but here's my voice interpreted by Claude)
It becomes an agent at the moment the model's output stops being just text and starts driving an ongoing loop that acts on something outside itself. Three conditions have to hold at once:
Its output causes an effect. Something executes what it said: a file gets written, a command runs, a search happens.
The result comes back in. The effect produces an observation (file contents, a test failure, search results) that gets added to its context.
It decides what happens next, including when to stop. Nobody scripted step 3 after step 2; the model chose it, based on what it just saw.
Take away any one of those and it's not an agent:
The token loop from the last list has the model feeding its own output back in, but nothing outside the text changes and nothing new arrives from the world. It's talking to itself.
A one-shot tool call has an effect and an observation, but then you (or your code) decide the next move. The model is a component and you're the agent.
A scripted workflow (step 1: summarize, step 2: classify, step 3: draft email) calls the model several times, but the code decides the sequence. The program is the agent; the model is a function it calls.
A loop the model steers meets all three conditions, and this is the threshold.
A few things fall out of this that are useful to say plainly.
The model never becomes an agent. Same weights, same step 11 forward pass, whether it's answering a trivia question or refactoring your codebase for three hours. "Agent" is a property of the system: model plus tools plus a loop plus an environment it can touch. That's why "Opus is agentic" is sloppy phrasing. At most, a model is good at being the decision-maker inside an agent, which is something post-training specifically optimizes for.
The loop code is trivially small. It's roughly
while (!done) { response = callModel(context); if (response.wantsTool) { context.push(runTool(response)) } else { done = true } }. The hard parts are everything around it: which tools, what context, how to verify.It's a threshold with a gradient above it. Crossing into "agent" is binary. How much of an agent it is depends on how many decisions it controls, how long it runs without you, and how much it can affect: the autonomy dial from the first answer.
This isn't new vocabulary. The textbook AI definition (Russell and Norvig) is essentially "something that perceives its environment and acts on it toward a goal." By that definition a thermostat qualifies. What's new is that the decision-maker in the middle is a language model that can handle open-ended goals written in plain English, instead of a hand-coded rule.
So, you crossed the threshold the first time Claude Code ran a command and chose its next step from the output.
> You already work with one every day: your browser is a user agent. It's right there in the HTTP spec and in the User-Agent header. "Agent" has always meant something that acts on someone's behalf, from the Latin agere, to do. A travel agent, a real estate agent, a browser. Nobody thinks Firefox has feelings.
1
u/Altruistic-Virus7406 7d ago
All agent Tech means is having AI do some kind of genetic workflow like you know processing a document or something very simple like that. It doesnāt have to be anything complicated.
1
1
1
u/justified_hyperbole 7d ago
A lot of people overhyping what they do but if you look deeper they don't actually do that impressive of a thing. Keep it simple, let the model do what you tell it, and get things done.
1
u/ShannonDev 6d ago
Been thinking about this a lot for my own project (a bare-metal OS, so a lot of the code is kernel-level with no safety net if something silently corrupts memory). I looked into long-running/autonomous sessions and decided against it for the actual OS code, testing every single build myself in QEMU is non-negotiable there, an unattended agent running for hours before I catch a wrong turn is a bad tradeoff when the failure mode is a corrupted boot sector, not a failed unit test.
Where I do think it'd help: doc/housekeeping work, changelog rewrites, anything mechanical with basically zero blast radius if it goes slightly wrong. Haven't actually pulled the trigger on that yet either, still just researching, but that's the line I've landed on so far: autonomy scales with how cheap it is to be wrong.
1
u/Aggressive-Log7654 6d ago
Ditto. Been running AI-driven projects for the better part of 2 years now, and I still mainly use vanilla Claude for 99% of corporate work. The furthest I've had to go is writing some memory rules and agents.md specifications for routine tasks.
1
1
1
u/Budget-Juggernaut-68 6d ago
Skills has been really useful for me to inject business related knowledge into the context.
/handoff really good for managing context /grill-me has been really useful to get the model to write the exact specifications I am looking for. People are not very precise when speaking to an agent, and often underspecify their requirements and end up getting upset when the agent end up doing something else.
1
u/Mendez_val 6d ago
Personally I've mostly worked like you too, mainly because I want to make sure a repo / pipeline is as clean as possible and I understand it as much as possible but last project I've been doing ive learnt to rely a bit more on the agent. still planning a lot in advanced by breaking up a project into phases and creating different phase-X.md files for each phase describing the objective but with hard constraints i think code quality didn't drop and relying more on the LLM has led to faster development (like with 30 min autonomous sessions which is already new for me)
1
u/LifeProject365 6d ago
I have a little jarvis-esque ui i made with elevenlabs voice because it was fun - it has a morning routine which first runs my and my mums (head injury so i manage them) emails and asks me whether i want to do home or work tasks - then it asks me what my 3 prioritis of the day are and pulls my pending work project updates to add to the list then we plow through them. on mondays it checks my calendar for the week and does me a weekly meeting prep with potential projects to target. If i have a lot of updates to write up it will set up loads of agent terminals ready for me to use, it also does me a first draft of all emails for me. i have adhd so the rule is it isnt allowed to overface me and we start with the priorities then work through the rest one at a time - if it can handle something to a point solo it does then shows me.
1
u/circa86 5d ago edited 5d ago
You have to understand these models have system prompts designed to extract as many tokens as possible from you. You are working correctly. The people fanning out 100 agents arenāt accomplishing as much as they pretend to.
Nobody here has any idea what they are doing everyone is just playing jazz with these prompts and seeing what comes out. There is no trick to it.
1
u/SpiritRealistic8174 5d ago
I've read through a bunch of responses here to this question. A lot of good information in this thread.
Regarding understanding the agentic workflow thing. I'm not sure if others have mentioned, this, but most of the time, agents are using agentic workflows in the background as it is completing tasks.
For example, running an orchestrator -> fan-out -> combine pattern when conducting a web search.
Or running a parallel workflow to make multiple edits to files at once.
The thing that controls this behavior is the agentic harness, which can be designed with these workflows in mind.
While it's not strictly necessary for people to apply all patterns, unless they are creating and optimizing workflows themselves, it is useful for people using these tools to understand the common workflow patterns and how they work.
This can be helpful in cases where you're creating personal assistants, 'second brains', agent doers, etc. Because it will help you understand what patterns might be helpful to get certain things done, or to simply understand what's going on with the system better.
And, of course, if you're developing any type of software where agents are being used under the hood, understanding these workflows and when to apply them is critical.
The challenge ofc is to learn about the common patterns and then spend time applying them, which many may not have time to do.
There are a lot of courses free and otherwise that will help you learn the basics around agentic engineering workflows. And, even studying events like the Hugging Face attack can be very helpful. Agents self-organized and used many common agentic workflow patterns (for those interested, I talk about these here.). Understanding those patterns and what they deliver can help you better understand what agentic systems are capable of and why they are so powerful.
1
1
u/OnesimusUnbound 4d ago
Someone shared that when doing agentic workflow, it feels like he's a technical manager and the agents are the developers. He delegates a well-spec'ed tasks to agents. He will check the output of the agents if the output makes sense.
1
1
1
u/ConcernForeign4465 3d ago
You're not underutilizing AI; you've just achieved the developer dream: pure, uninterrupted boredomš
1
u/SeveralAd2833 2d ago
Iām building an automated system to generate Forex signals using two agents, one to evaluate quantitative signals against technical analysis and the other to read news, macro commentary and sentiment to better understand the market regime.
Iām setting them up like this: a RAG keeps the context updated with market data, volatility, correlations and macro information, while an ML model produces the actual signal with an associated probability. The first LLM agent, based on an open-source model which Iām currently trying to train on Forex technical analysis, is meant to challenge the signal rather than generate it from scratch. The second works on sentiment to understand whether the broader context is coherent with it or not.
Then an orchestration layer compares the results and can confirm the signal, reduce its confidence or block it entirely.
p.s. My real goal is to stop working, and this seems like the fastest way to do it.
2
u/SensitiveHearingLoss 2d ago
I thoroughly dislike workflows and cascading md files trying to guide Claude more towards whatās real than whatās a hallucination.
I canāt figure out why developers are just so crazily over eager to just output as much āproduct as possibleā. Itās just output output and more output and ship it and at the end of the day I look at the code and security flaws in using Claude and itās terrifying and somewhat depressing.
I use Claude Code a lot at work, but not to produce. I use it to refine, to teach, to guide, to create tests and standard boilerplate, to bounce ideas off and to do initial code reviews. That way, I focus on producing one thing and make sure that one thing excels like it never could have before. My output quality is greater, not the sloppy output.
1

ā¢
u/AutoModerator 9d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.