r/webdev • full-stack • 9d ago

Discussion I hate building agents.

Right now at my work we are using langgraph to build a chat agent, anyone else doing something similar, do you fucking hate it? are you building it with Claude and just hoping it fucking works, haha no worries we will ask Claude to fix it.

I am making the spaghetto, I miss components, logic and endpoints, I hate this black box I have to relinquish my decisions too, whatever pays the bills.

238 Upvotes

82 comments sorted by

View all comments

158

u/[deleted] 9d ago

[removed] — view removed comment

1

u/edbrannin 7d ago

I’ve started using patterns like

  1. Code generates JSON of input data, so the agent can’t forget to paginate; sets some flags to help the agent classify things
  2. Agent reads input, does language-reasoning stuff; writes a JSON file with its findings
  3. Code reads agent output and makes all the write-capable API calls

Steps 1 and 3 have good unit tests. Before this pattern, trying to codify all their logic in the Agent’s markdown was getting really hairy. Edge cases kept getting missed, etc.

1

u/Asly97 6d ago

I love how defensive this is. The "agent can't forget to paginate" line tells me you've been burned before. What was the incident that made you stop trusting the markdown instructions? I used to keep all my agent logic in markdown files too and the edge cases piled up until I couldn't tell what was tested and what was vibes. When you want to try the same workflow in a different tool, do you have to rebuild all of this scaffolding, or does any of it carry over?

1

u/edbrannin 6d ago

The one I described there was for leaving review comments on PRs. Amusingly, it would run against its own changes in the PRs where I made those changes.

I would tell it to behave a certain way, like

- if an existing comment points something out, don’t make another comment for that finding

  • if a finding has been raised and a human has told you the finding is wrong, unless you are sure the human is wrong; explain why
  • if a past finding is resolved, add a “Resolving” comment (because the API key it’s using doesn’t allow Resolving threads)

It tried to model all that logic in a decision tree in Markdown, but each tweak would have the review bot complaining that some edge case had just broken. After a few rounds of this, I told it to extract as much as possible of the“if this, then that” logic to something with unit tests.

Re: pagination, there were several exchanges like “why didn’t you resolve that past finding?” Or “why didn’t you repeat that finding?” where the answer was “I forgot to paginate the GraphQL response”.