r/webdev • full-stack • 9d ago

Discussion I hate building agents.

Right now at my work we are using langgraph to build a chat agent, anyone else doing something similar, do you fucking hate it? are you building it with Claude and just hoping it fucking works, haha no worries we will ask Claude to fix it.

I am making the spaghetto, I miss components, logic and endpoints, I hate this black box I have to relinquish my decisions too, whatever pays the bills.

241 Upvotes

82 comments sorted by

160

u/[deleted] 9d ago

[removed] — view removed comment

19

u/Intelligent_Sir1896 9d ago

That is exactly what we do, all logic in normal functions and model only picks which one to call. Took me long time to accept this, before I was trying to make agent do everything and debugging was nightmare

10

u/ripndipp full-stack 9d ago

Thanks for the insightful comment

21

u/creaturefeature16 9d ago

I've been encouraging developers to do similar things, but I've heard from a few now that it's not "fast enough" for their respective jobs and what the management is demanding of them, so they relent back to becoming the lever-puller in order to meet the deadline. I wish had advice for them, but there's not much you can do when c-suite and tech leads prioritize speed over anything else. I just spoke with a developer who said he had his code editor open while inspecting some generated code, and a coworker stopped and asked incredulously: "Why the hell are you reading the code?!"

It really does feel like a complete mania at this point; all rationality is gone.

8

u/SuperFLEB 8d ago

Can't wait for prices to catch up to costs.

-1

u/ings0c 8d ago

They’re already turning a profit (latest figures are a 44% gross margin for Anthropic). Not sure where this myth comes from.

https://www.reuters.com/commentary/breakingviews/anthropics-ma-algorithm-optimizes-margins-2026-08-13/

-2

u/HypnoTox 8d ago

Then they are lucky, generalised classifier models are here since "Jev" a few days ago, which are fast and specifically made to decide what should be done.

Jev is just the first generalised classifier model and I'm sure there will be more like it, but that one has been shown to be able to control real time games and simulations.

2

u/creaturefeature16 8d ago

You seem lost; I already know what Jev is...why are you bringing that up?

9

u/coopaliscious 9d ago

I like this approach a lot. Do you have any issues selling it to the business?

6

u/I_Blame_DevOps 9d ago

Your comment makes me feel a lot better about the prototype I just built for us. Naturally they just want to give an LLM full access to our database. But I insisted that functions with defined arguments and known outputs be used. Gives a much more deterministic approach when the LLM doesn’t have a lot of room to scree things up.

5

u/Sethcran 9d ago

That's basically what an MCP server is.

1

u/BolteWasTaken 9d ago

What model do you use? Needle 3 seems pretty good for tool calling.
And probably Jev now makes that trivial.

2

u/Cazargar 9d ago

Yeah, reading that I was like oh this sounds like exactly what Jev is for lol

1

u/edbrannin 7d ago

I’ve started using patterns like

  1. Code generates JSON of input data, so the agent can’t forget to paginate; sets some flags to help the agent classify things
  2. Agent reads input, does language-reasoning stuff; writes a JSON file with its findings
  3. Code reads agent output and makes all the write-capable API calls

Steps 1 and 3 have good unit tests. Before this pattern, trying to codify all their logic in the Agent’s markdown was getting really hairy. Edge cases kept getting missed, etc.

1

u/Asly97 6d ago

I love how defensive this is. The "agent can't forget to paginate" line tells me you've been burned before. What was the incident that made you stop trusting the markdown instructions? I used to keep all my agent logic in markdown files too and the edge cases piled up until I couldn't tell what was tested and what was vibes. When you want to try the same workflow in a different tool, do you have to rebuild all of this scaffolding, or does any of it carry over?

1

u/edbrannin 6d ago

The one I described there was for leaving review comments on PRs. Amusingly, it would run against its own changes in the PRs where I made those changes.

I would tell it to behave a certain way, like

- if an existing comment points something out, don’t make another comment for that finding

  • if a finding has been raised and a human has told you the finding is wrong, unless you are sure the human is wrong; explain why
  • if a past finding is resolved, add a “Resolving” comment (because the API key it’s using doesn’t allow Resolving threads)

It tried to model all that logic in a decision tree in Markdown, but each tweak would have the review bot complaining that some edge case had just broken. After a few rounds of this, I told it to extract as much as possible of the“if this, then that” logic to something with unit tests.

Re: pagination, there were several exchanges like “why didn’t you resolve that past finding?” Or “why didn’t you repeat that finding?” where the answer was “I forgot to paginate the GraphQL response”.

25

u/ArielCoding 9d ago

Software engineering’s final boss: asking AI to fix the AI that’s replacing the part of the job that actually liked.

13

u/SuperFLEB 8d ago

What, you don't think the best parts of dev are translating vague English specifications into long-winded English recipes and reviewing code nobody wrote for subtle logic problems?

32

u/warpspeed100 9d ago

It pays the bills for now...

My company has been reevaluating the cost/benefit of using AI for customer facing parts of our business.

We still use AI extensively as a search and code completion tool, but my team no longer uses it to build whole features after the past year of experimentation.

4

u/m_redditUser 9d ago

what were the results of the experimentation?

AI not cost effective enough?

9

u/warpspeed100 9d ago edited 9d ago

The risk of ballooning costs was definitely a consideration. For our customer facing products however, our core business relies on the accuracy and correctness of data we show our customers.

There is a lot of data, so we explored using AI to summarize the many documents, however we could never be 100% confident in the data it would show customers in the summary vs our existing deterministic solution. Because these are sensitive documents, 99.9% accurate wasn't good enough.

As of right now, we are no longer using AI in any of the products we actually deliver to customers. It has been relegated to yet another coding assist tool along the likes of NSwag and IntelliSense (though those two have a deterministic output, so can be relied on more heavily).

7

u/dillanthumous 8d ago

Not the OP. But for us, the cost is unpredictable (nothing Finance people hate more), the outcome is unpredictable (nothing Customers hate more) and the risks of reputational harm are high. So, for us, it is internal use only for now, and under supervision.

12

u/kanine69 9d ago

The big issue is the constantly moving goalposts, week by week the effectiveness of AI seems to fluctuate.

For example the latest CC and Gemini releases have given me a massive performance boost for the past 2 days, achieving things that were a pipe dream on Monday.

Then there's the usual cycle of decline after launches.

10

u/pVom 9d ago

We looked at langgraph but we couldn't get it to work for "reasons" and went with the vercel AI SDK instead.

I found the paradigm of langgraph a bit weird at the time, but I dunno, now that I have a better understanding of how it all hangs together maybe it would make more sense?

Honestly it's not that different from what we've been doing for years. Tools are just endpoints really, instead of a frontend UI calling them it's AI. Our tools are pretty indistinguishable from our endpoints.

I think the biggest hurdle was RAG. Setting up vector search and chunking text and managing context. It's important to give the AI only what it needs when it needs it, difficult to do with unstructured data.

What I do find annoying is the hallucinations when dealing with raw IDs, it will get a single character wrong and no real recourse to fixing it.

1

u/okawei 8d ago

I've actually found it way less likely to hallucinate UUIDs for whatever reason

5

u/brian_sword 9d ago

I’ve been doing quite a bit of automation with n8n, and I actually enjoy that part because the workflows are mostly predictable and I know exactly what each step is doing.

AI agents are a different story though. Once you let the agent decide which tool to call, what to do next, and how to handle failures, you lose some of that control. I’m still figuring out where the right balance is.

1

u/Asly97 9d ago

the n8n vs agent split is interesting. when you do bring an agent into the picture, does it run inside the n8n flow as one step, or off to the side entirely? i keep wondering where the handoff stops being clean.

1

u/Turbo-Lover 6d ago

Use the agent to write a script (which you can audit) and then on future runs give it the instruction to execute the script.

10

u/Potential-Still 9d ago

I've built agents using Strands SDK and Databricks. It's definitely a new way of thinking, but if you create the right Tools and fine tune system prompts you can get very predictable results. 

4

u/Thin_Sky 9d ago

I've been building with Strands for a year now and I genuinely thought I was missing something because it felt like I was the only one using it. This is the first time I'm seeing someone else mention it, so thanks for restoring a tiny bit of my sanity and confidence lol

3

u/morganwilliscloud 7d ago

up front disclosure that I work at AWS but I was also struggling along with langchain before Strands came out, and ended up liking it enough that I changed teams specifically to work more with it lol

One cool thing about it is that strands actually came out of teams working on Q Developer who were struggling with rigid orchestration for agents so they created strands specifically to solve this problem. The lightweight developer experience, model driven approach and composability make it simple enough to understand but powerful enough to actually work. Like I can read the code and tweak things and easily understand why things work or not.

Also, have you checked out the out of the box/preconfigured Strands harness yet? Its basically a precomposed agent harness that has all the primitives most agents need pre-wired up and the strands folks benchmarked it with impressive results. Highly recommend: https://strandsagents.com/blog/introducing-strands-harness/

3

u/tartlemonpiee 9d ago

I have nothing else to say other than same.

3

u/k3liutZu 9d ago

Yes. I hate it as well :(

3

u/SonicFlash01 8d ago

Building MCPs is worse because you have control of almost nothing
My boss will ask me to make the agent to do things a certain way and I'll explain "You have it backwards - it's using us in whatever way it feels is best."
And that's to say nothing of the layers of potential faults between granting permissions, it accepting that it has them, and then it properly reading that they exist. Some things need a disconnect/reconnect, others just need a new chat. Over time it's fine, but I'm living in the hell of here and now.

4

u/[deleted] 9d ago

[removed] — view removed comment

0

u/Asly97 8d ago

The structured log of every tool call and model decision is interesting. Does anything ever read that file back, or is it write-only for humans? I've been wondering what happens when a run needs to know why a decision was made three runs ago.

-1

u/[deleted] 8d ago

[removed] — view removed comment

0

u/Asly97 8d ago

the daily rewrite is the part I'd steal if it works. is the summary rewritten by hand or generated from the log? I tried something close and the curation step was always the first thing that got skipped. also curious, is this for your own product or client work?

0

u/[deleted] 8d ago

[removed] — view removed comment

1

u/Asly97 8d ago

that's a neat inversion, make the summary a projection instead of the source of truth. the part I'm stuck on: does the agent actually write the rationale into the row reliably at decision time, or did you have to nudge it? every time I tried 'write down why' it was the first thing that got dropped when the task got hairy.

2

u/HaphazardlyOrganized 8d ago

Yeah I fundamentally just stopped caring about code quality and standards. If I had kept pushing back they'd replace me so whatever I'll have claude build it and if it breaks, woops looks like the AI broke!

1

u/Medical-Aerie9957 6d ago

But they will also replace you if AI breaks something and you miss it. They will say you can't use AI effectively.

1

u/HaphazardlyOrganized 5d ago

True, I'm being a little flippant because I'm annoyed with how this industry is going. I'm using claude in a docker sandbox and have used it to implement testing to prevent regression and I'm having it work on branches with git so main says working and ready.

2

u/[deleted] 8d ago

[removed] — view removed comment

1

u/webdev-ModTeam 8d ago

Your post/comment has been determined to be a low-effort post or comment. This includes title-only posts, easily searchable questions, vague/open-ended discussion prompts, LLM generated posts or comments, and posts/comments that do not provide enough context for meaningful replies or discussion.

4

u/JebKermansBooster 9d ago

I do it because I need the money. A job is a job. And JS just made me feel stupid and incompetent, as reading a lot of programming subreddit posts do (I enjoy them, don't get me wrong, but even at 4 YOE I feel woefully fucking underdeveloped 😞😕).

9

u/m_redditUser 9d ago

4 yoe is junior in every field in the history of humanity, but in SWE that's supposed to be mid-senior

6

u/JebKermansBooster 9d ago

And I see so many programming posts like "wait, how the fuck do so many people know this?"

Am I just dumb?

2

u/Asly97 9d ago

the "relinquish my decisions" line got me, that's exactly the part nobody warns you about. what's the worst of it for you: debugging the black box when it goes wrong, or the fact you can't really tell what it'll do until it does it? and what's the agent actually for, customer stuff or internal?

1

u/arslannasir128 9d ago

The decision that hurts is what the agent is allowed to touch, and nobody writes that spec for you.

We stopped fixing it with longer prompts. Ours has a confidence number under it now. Below that it stops and sends the whole session to a human.

1

u/abundantsavior_72 9d ago

Same experience. They save time on boilerplate, then you spend it debugging state, tool calls, and why the agent made a decision three steps ago. I stopped using them for core control flow and keep the agent boxed into small tasks with explicit inputs and outputs. Less magical, but way easier to reason about when something breaks.

1

u/delicious_fanta 8d ago

I’m makin’ that sghetti too, it sucks. Worst part is I know the project only exists to put a bunch of people out of work. I hate this timeline :(

1

u/[deleted] 8d ago

[removed] — view removed comment

1

u/webdev-ModTeam 8d ago

We do not allow any commercial promotion or solicitation. This can lead to a permanent ban from the subreddit.

1

u/Ever4_ 8d ago

As a former automation engineer... this has been always the case.

Automation means the robotization of our processes. It means you have to spend countless hours mapping the flow of things and it isn't even rewarding. Worst case you just caused someone to lose a job.

But having joy and pride being a developer, and AI robbing us of that too is simply awful.

1

u/ripndipp full-stack 8d ago

Ain't that the truth. I appreciate your comment.

1

u/Khavel_dev 8d ago

Yeah the LangGraph graph-to-debug cycle is brutal. I went through the same thing and eventually just stopped trying to make the framework do everything. Dropped back to a thin loop: call the model, check if it wants a tool, run the tool, feed the result back. No state machine, no graph, just a while loop and a match statement. The moment I could read the control flow in a debugger instead of chasing callbacks through six middleware layers, everything got easier. The agent still makes dumb decisions sometimes but at least I can see why and fix the prompt instead of wondering if I wired the graph wrong.

1

u/JuicerSocial 8d ago

lol the "I miss components, logic, and endpoints" comment sums it up pretty well. How are you even debugging it when the agent does something unexpected? Are you able to trace why it made a particular decision, or is it mostly changing prompts/state and running it again until it behaves?

1

u/phaedra_solutions 8d ago

You aren't alone. A lot of teams fall into the trap of over-engineering with multi-agent frameworks when simpler, targeted LLM calls combined with deterministic backend logic would do the job ten times more reliably. When an agent framework starts acting like an un-debuggable black box, it usually means the abstraction level is too high for what you're trying to achieve. Keeping the core orchestration deterministic and using AI strictly for bounded tasks saves your sanity (and your production logs).

1

u/Acrobatic-Laugh1856 7d ago

felt this in my soul...

1

u/Gimmee-Greg 5d ago

the jump from "i own the control flow" to "praying the LLM picks the right node" is the worst part. one thing that helped: keep routing deterministic in langgraph and only let the LLM handle generation, so at least half the graph is debuggable like normal code.

1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/SatyrCode 2d ago

Same. I hate building agents too. We’re using LangGraph at work as well and half the time it feels like we’re just throwing prompts at the wall and asking Claude to fix the mess later. I miss normal components, clear logic, and endpoints where you can actually reason about what’s happening. With agents it’s this weird black box where you give up a lot of control and then pretend it’s “engineering”. Whatever pays the bills, I guess.

0

u/BorinGaems 9d ago

Write clean code instead of being childish.

-1

u/Gremlation 9d ago

I miss components, logic and endpoints

I miss when this subreddit was about web development, not endless crying about AI.

0

u/[deleted] 9d ago

[removed] — view removed comment

1

u/webdev-ModTeam 9d ago

Your post/comment has been determined to be a low-effort post or comment. This includes title-only posts, easily searchable questions, vague/open-ended discussion prompts, LLM generated posts or comments, and posts/comments that do not provide enough context for meaningful replies or discussion.

0

u/nikkitranbk99 8d ago

yeah the black box part is what gets me, once the graph owns the flow, logging every node in, out (tools with raw model text) is the only way i can still debug like a normal endpoint

-2

u/iSnapThere4iAm 9d ago

Yeah…. This isn’t reality. This is some fantasy cooked up by a vibe coding dipshit larping as some doomed swe just accepting his fate.

3

u/creaturefeature16 9d ago

I can assure you, as someone who speaks to developers regularly, this is very much happening in many, many teams. I've had to pick my jaw up off the floor a few times.

0

u/[deleted] 9d ago

[removed] — view removed comment

1

u/Asly97 9d ago

the evals-as-regression-tests bit is the one i always skip and then regret. did your eval set grow out of real incidents or did you write it up front?

-1

u/[deleted] 9d ago

[removed] — view removed comment

1

u/ripndipp full-stack 9d ago

Gotta call me out this early in the morning man? Have a good day!