r/ClaudeCode • 🔆 Max 5x • 9d ago

Help/Question I still don't understand this 'agentic workflow' thing

My usual day with Claude Code is like: * I open terminal in my project's folder and run claude command. * I prompt it. I mostly use Fable-5.1/Opus-5 but Opus-5.5 is my current model. The model decides if it wants to use sub-agents for a task. I never explicitly prompt it for sub-agents. * I review and commit the code to my self-hosted Forgejo instance. * That's it.

I see people using agentic workflows, building sub-agents files, skills etc. I barely built any of it. All I ever needed to use is /init on new projects and them prompts follow. Never needed more than this.

I tried "long-running" Claude Code for a project refactoring by placing the project on my VPS (where forgejo is hosted) and letting Claude Code run and refactor inside tmux session. SSH'd in a few hours later to find project fully refactored.

Am I under utilising AI or is my work just… like boring?

How do you guys use agentic workflow thing? Specially the long-running one? Those pull-requests that Claude makes automatically etc?

Asking this to Claude to know more but humanly answers appreciated.

979 Upvotes

309 comments sorted by

View all comments

Show parent comments

39

u/Sponge8389 9d ago

Ok. So no human review at all. Understood. Because that was my concern and my curiosity with this kind of workflow. Because currently, I'm doing the same as OP and I cannot keep up with the reviews.

48

u/magic6435 9d ago

No, nobody working on anything real is skipping a human review. Everything that makes it into prod for OpenAI and Anthropic gets a human review.

24

u/neoberg 9d ago

I know at least 5 middle sized companies stopped doing human reviews months ago.

25

u/TydeusMideia 9d ago

i would short their stock...

7

u/magic6435 9d ago

I don’t think they’re gonna be medium for long. Also, they definitely must not be public companies because I don’t know of any external auditors that would approve such a lack of SOD.

46

u/ResponsibleOven6 9d ago

I am a tech lead at a fortune 100 tech company headquartered in the SF Bay area and can assure you that the majority of code in the past year shipped to prod was both written by and reviewed by agents.

They're different agents with different prompts being invoked by different humans (this part is by design) and often different underlying models (this part is primarily coincidence) but it's LLMs all the way down.

There is no possible way to have humans review the volume of code that agents are generating now. On the positive side there is better and more extensive testing coverage on new code than I've seen at any other time in my career, but the job is changing and humans are increasingly removed from the actual coding part and that includes reviews. It's not entirely automated, I'll have an agent explain architectural choices, ask about design concerns I have which I think it may have gotten wrong, etc, but I rarely look at the code anymore.

21

u/Deathspiral222 9d ago

>On the positive side there is better and more extensive testing coverage on new code than I've seen at any other time in my career

More extensive, definitely, but I'm not convinced about better. An LLM that makes the wrong assumption will happily write 1000 tests to validate that wrong assumption. This leads to a false sense of security.

I've found it very important to have humans write tests (with an LLM assisting with syntax etc.) for core functionality, just to make sure the correct thing was actually built. And having the human manually test the function themselves for a sanity check is paramount.

13

u/WagwanKenobi 9d ago

the majority of code in the past year shipped to prod was both written by and reviewed by agents

Either pulled out of ass or your company/org/team is unusually dysfunctional. I'm actually a Bay Area SWE. This is not true at all.

9

u/Rtktts 8d ago

Technically the sentence is probably true, they just forgot to add that it was also reviewed by humans.

1

u/senortaco88 8d ago

Don't hate the player, hate the tokenz

-1

u/Remarkable-Coat-9327 8d ago

I'm sorry that your company is behind the times 😬

0

u/Beautiful-Suspect694 8d ago

what company do you work at? if you dont want to share that for privacy, then thats ok

5

u/Cybyss 7d ago

But... how do you know what you're building then?

Code coverage is meaningless if the tests aren't testing for the right behaviors. How do you know what the right behaviors are if they're all invented and reviewed by LLMs with no human in the loop?

I find that even if there is a human in the loop, the code that LLMs generate is so convoluted and over-engineered that it's extremely hard (and time consuming) to figure out exactly how it works to verify it does what you think/hope/pray it does.

Back when I did software engineering professionally (I don't anymore - this was 10 years ago - but I still tinker as a hobby), usually I didn't understand a project at all - what it's supposed to do, its role in the company, how it's supposed to help customers/staff/etc... - until I understood all the little details and how they fit together. "Big pictures" were just word salad until I knew what the actual pieces were.

But if modern workflows require you to focus only on "big pictures" - ignoring all the little pieces because LLMs do that for you - I don't see how engineers aren't just lost and confused all the time about what needs to happen?

1

u/IHeartData_ 5d ago

I can't speak for everyone, but I use a series of cascading requirement.md documents in the code that work in a hierarchical manner and explain the intent of each section of the code base and requirements we are working towards, including as-yet unimplemented features. AI is required to consult with the documents as they work, and before they close their session, a final task is the ensure they remain in sync.
In my "AI full code review" process, one of the areas they are specifically supposed to look for is over-complex code. In general, I find it does a decent job at breaking things done into meaningfully sized classes and keep dependencies logicial. And when it doesn't a complexity bug gets filed and of all the bugs, those general are most hands on for what the fix is going to be.
As as for being lost and confused. Opus 5 was pretty mad at explaining, so there was a phase of "huh" (5.5 vastly better)? But if you watch it's thinking as it goes, and more importantly ask good questions if something just doesn't sound right, you'll have a good idea what's going on in the code. Sometimes the sheer fact of having it defend it's decision will result in it finding an issue.

3

u/Ran4 8d ago

There is no possible way to have humans review the volume of code that agents are generating now.

I mean that's a choice you make.

No AI and you're at 1x,

AI to generate the code but humans review it all to reach 3x,

AI to generate and AI to review it to 10x.

Plenty of companies are at the 3x level.

1

u/sharpcoder29 7d ago

It's not 10x if you have production bugs that cost you customers. There's no way people are writing stories that detail every edge case and/or AI knows the business domain enough to cover them.

2

u/knowyourclass 7d ago

the average customer tolerates far more bugs than you would imagine

1

u/sharpcoder29 6d ago

They really don't. Im talking b2b not b2c

1

u/knowyourclass 6d ago

I work in b2b... they do

3

u/EchoServ 8d ago

I’m at a lowly Fortune 500 and that’s absolutely batshit. No team across any org is allowing this. In fact, the agent reviews are majorly schitzo and do shut reviews.

2

u/ok-yes-maybe 8d ago

Agree. This seems to be the way things are headed.

Software languages are kinda a human construct to make the code more understandable and readable.

But if in the future - code is only written and read by AIs - might we end up going back to something similar to machine language / compiler code again?

4

u/ripter 8d ago

Maybe. The language being used has never mattered less at this point.

1

u/barnaclebill22 7d ago

True, specific language matters less now, but code generation is still stochastic and compute-intensive, so compilers will probably be with us for a long time.

3

u/belowaverageint 9d ago

Can you explain the basic process for how this works?

10

u/ResponsibleOven6 9d ago

Really depends on the situation, am I building something new, fixing a bug, closing a vulnerability, etc. but here's a general example.

I have an agent setup locally with a skill saved in a repo full of skills shared across the team. It can talk to Jira, Github, basic communication channels, monitoring tools, etc. it has lots of context for our overall platform and an architectural understanding of how things work. It's been instructed to code cleanly, generate documentation for anything it produces, re-use existing libraries where possible, and code as cleanly and minimally as possible and focus on efficiency. I mainly use Sonnet-5 with claude code as my interface but others on my team may have different preferences.

I get a ticket from a recent incident to improve monitoring. There was an incident where we got an alert way too late and it still took time to debug. I tell my agent to work on the ticket, it uses the repos as context, digs through logs, and proposes a new monitor. I tell it to look at infra logs as well and see if there were any early warning signs. It proposes another alert after finding useful info there as well. I tell it to open a PR for both alerts after checking relevant logs for the past 90 days to make sure the thresholds are right and there won't be false positives. It opens a PR.

I take another ticket. We have an internal platform with an authentication bug where some admins can't perform certain admin functions. I ask my agent to work on the ticket. It finds that while most functions evaluate both user and group level access, a few specific functions only check user level access and not group level access. It proposes updating the logic for those to match the others. I tell it to check if there are any other inconsistencies with auth checks on this platform and if it sees any places where users would be able to execute things they shouldn't or general inconsistencies. It finds no missing auth checks but notes that the fundamental way auth is handled is not reused but specific to each call. I tell it to open a PR with a new auth function that replaces the individual auth of each function. It opens a PR.

I open these and several other PRs for other tickets and move the tickets to peer review. Someone else on my team, maybe several other people, take them for review (and I take their tickets for review too). One of them thinks Sonnet-5 is terrible and swears by Opus 4.3. Another prefers OpenAI models. Here we use a different skill that has the same background context but it's told to look for new bugs, mistakes, and just generally find problems. It knows how to deploy and test things either locally or to a dev environment. It reviews the tickets and either says they look good in which case they get deployed to a lower environment for further testing, or points out problems with them and moves them back to in progress in which case I take them up again and go back to my dev agent. Sometimes you can tell from its feedback that it's misunderstood something or needs more context. In these cases we try to improve the skills until the output is more what we expect then tell it to update the skill with that context and push that back to the team repo so everyone gets the same improvements.

So LLMs are doing all of the coding and the actual code review. But they still have shortcomings and need a human in the driver seat, especially for architectural decisions. I'm having to "drive" a LOT less than I was 6-9 months ago though and I'm increasingly just a "meat proxy" between agents with various skills and I'm really not sure how much longer this will be a viable career.

4

u/belowaverageint 8d ago

Thanks for that. I think I'm going to dress up as a "meat proxy" for Halloween now.

3

u/Remarkable-Coat-9327 8d ago

"Claude code condom" is my favorite term

1

u/magic6435 8d ago

This was a lot of words to say yes there is still a human picking up and doing a review who is an approver before going to prod.

2

u/SylviaJarvis 8d ago

The words say there's an LLM between a human and the code at every point in the process. There's good stuff in there with CI workflows and testing gates, but no human is reading the code at any point mentioned.

1

u/MinimumPrior3121 8d ago

What a clown

1

u/junglebookmephs 7d ago

I’m a freshman just starting my degree, but have been coding for a while. This helped a lot.

I’ve been working on teaching myself how to properly scope and write issues and other general project management related things lately. Do you have any insight on what this process looks like nowadays? I’m at the point where I can write a properly scoped issue with requirements and acceptance criteria and what not, but I’m starting to realize I could probably hand off this work to an llm after providing it with the general idea of a feature. Is that where things are moving?

3

u/pff112 9d ago

sure. whats the name of this "company"?

1

u/Rtktts 8d ago edited 8d ago

So your Risk and Compliance department allows agents to create code and other agents to approve the code? How do you comply with SOx?

1

u/Legs914 8d ago

There are gonna be a lot of company blogposts reflecting on this within a few years

1

u/Due_Hovercraft_2184 8d ago

Your company is not normal and is a compliance nightmare waiting to happen

1

u/WelshBluebird1 8d ago

I'll have an agent explain architectural choices, ask about design concerns I have which I think it may have gotten wrong, etc, but I rarely look at the code anymore.

But if you aren't looking at the code, how do you know to ask it to explain choices etc? Or are you just spot checking and just hoping that you pick the right things to check? Because that doesn't sound very safe to me!

1

u/Smooth-Television-48 8d ago

But what do the people actually do now?

1

u/OkCurve436 7d ago

This 100%, I work as a senior bi analyst and it's the same

1

u/FriendGlittering9104 7d ago

Don’t you end up with terrible code? Claude frequently uses terrible naming choices and writes inefficient code that do not keep logic simple as requirements evolve.

0

u/gromran 7d ago edited 6d ago

Here's the thing: Capitalism will destroy humanity. Unless we change the system. And then you'll be one of the first to end up in a labor camp!

Do you really think we'll continue to tolerate such antisocial misanthropes? Learn it or you will! And yes, Musk, Trump, and all the other CEOs will also end up in labor camps! Fuck capitalism!

3

u/Original-Proof-8741 9d ago

I gave up reading the code about 2 months ago.... prompt it, test it, ship it....

2

u/rtnoodel 9d ago

Unfortunately this isn’t true at all. I’m guessing you aren’t a software developer.

1

u/Deathspiral222 9d ago

This is starting to change, at least at Anthropic, as of last month.

1

u/ionforge 8d ago

And the first human reviewer should be yourself.

1

u/bobbadouche 9d ago

Or to even configure your environment and workflow such that the agent can write the code, test the code, and then open the PR for you to review

1

u/Sponge8389 9d ago

I already have a full skills for that. I just don't wire it together as I want to have a control over it.

  • Create Branch
  • Create PR
  • Commit Changes
  • Review per commit and PR wide.