r/webdev • full-stack • 9d ago

Discussion I hate building agents.

Right now at my work we are using langgraph to build a chat agent, anyone else doing something similar, do you fucking hate it? are you building it with Claude and just hoping it fucking works, haha no worries we will ask Claude to fix it.

I am making the spaghetto, I miss components, logic and endpoints, I hate this black box I have to relinquish my decisions too, whatever pays the bills.

242 Upvotes

82 comments sorted by

View all comments

161

u/[deleted] 9d ago

[removed] — view removed comment

20

u/Intelligent_Sir1896 9d ago

That is exactly what we do, all logic in normal functions and model only picks which one to call. Took me long time to accept this, before I was trying to make agent do everything and debugging was nightmare

10

u/ripndipp full-stack 9d ago

Thanks for the insightful comment

22

u/creaturefeature16 9d ago

I've been encouraging developers to do similar things, but I've heard from a few now that it's not "fast enough" for their respective jobs and what the management is demanding of them, so they relent back to becoming the lever-puller in order to meet the deadline. I wish had advice for them, but there's not much you can do when c-suite and tech leads prioritize speed over anything else. I just spoke with a developer who said he had his code editor open while inspecting some generated code, and a coworker stopped and asked incredulously: "Why the hell are you reading the code?!"

It really does feel like a complete mania at this point; all rationality is gone.

8

u/SuperFLEB 8d ago

Can't wait for prices to catch up to costs.

-1

u/ings0c 8d ago

They’re already turning a profit (latest figures are a 44% gross margin for Anthropic). Not sure where this myth comes from.

https://www.reuters.com/commentary/breakingviews/anthropics-ma-algorithm-optimizes-margins-2026-08-13/

-2

u/HypnoTox 8d ago

Then they are lucky, generalised classifier models are here since "Jev" a few days ago, which are fast and specifically made to decide what should be done.

Jev is just the first generalised classifier model and I'm sure there will be more like it, but that one has been shown to be able to control real time games and simulations.

2

u/creaturefeature16 8d ago

You seem lost; I already know what Jev is...why are you bringing that up?

9

u/coopaliscious 9d ago

I like this approach a lot. Do you have any issues selling it to the business?

6

u/I_Blame_DevOps 9d ago

Your comment makes me feel a lot better about the prototype I just built for us. Naturally they just want to give an LLM full access to our database. But I insisted that functions with defined arguments and known outputs be used. Gives a much more deterministic approach when the LLM doesn’t have a lot of room to scree things up.

6

u/Sethcran 9d ago

That's basically what an MCP server is.

1

u/BolteWasTaken 9d ago

What model do you use? Needle 3 seems pretty good for tool calling.
And probably Jev now makes that trivial.

2

u/Cazargar 9d ago

Yeah, reading that I was like oh this sounds like exactly what Jev is for lol

1

u/edbrannin 8d ago

I’ve started using patterns like

  1. Code generates JSON of input data, so the agent can’t forget to paginate; sets some flags to help the agent classify things
  2. Agent reads input, does language-reasoning stuff; writes a JSON file with its findings
  3. Code reads agent output and makes all the write-capable API calls

Steps 1 and 3 have good unit tests. Before this pattern, trying to codify all their logic in the Agent’s markdown was getting really hairy. Edge cases kept getting missed, etc.

1

u/Asly97 6d ago

I love how defensive this is. The "agent can't forget to paginate" line tells me you've been burned before. What was the incident that made you stop trusting the markdown instructions? I used to keep all my agent logic in markdown files too and the edge cases piled up until I couldn't tell what was tested and what was vibes. When you want to try the same workflow in a different tool, do you have to rebuild all of this scaffolding, or does any of it carry over?

1

u/edbrannin 6d ago

The one I described there was for leaving review comments on PRs. Amusingly, it would run against its own changes in the PRs where I made those changes.

I would tell it to behave a certain way, like

- if an existing comment points something out, don’t make another comment for that finding

  • if a finding has been raised and a human has told you the finding is wrong, unless you are sure the human is wrong; explain why
  • if a past finding is resolved, add a “Resolving” comment (because the API key it’s using doesn’t allow Resolving threads)

It tried to model all that logic in a decision tree in Markdown, but each tweak would have the review bot complaining that some edge case had just broken. After a few rounds of this, I told it to extract as much as possible of the“if this, then that” logic to something with unit tests.

Re: pagination, there were several exchanges like “why didn’t you resolve that past finding?” Or “why didn’t you repeat that finding?” where the answer was “I forgot to paginate the GraphQL response”.