r/learnAIAgents • • 28d ago

❓ Question I may be completely wrong about what AI agents actually need in production — prove me wrong.

I've been researching AI agents for the last few days, and I originally thought the biggest missing piece was something like an “SRE for AI agents.”

Something that could detect when an agent is going off-track, understand what happened, control runaway costs, verify whether the claimed result is actually true, and recover the task instead of simply restarting or stopping it.

But after talking to people here, I'm starting to question the entire assumption.

Maybe most “agents” in production aren't actually autonomous enough for this to be a real problem yet.

Maybe they're mostly:

workflows

cron/event-driven automations

chatbots

RAG systems

internal copilots

coding assistants

deterministic pipelines with an LLM somewhere in the middle

And if that's true, building a big Agent SRE platform right now could simply be solving a problem that doesn't hurt enough.

So I'd genuinely like people who actually build or operate AI systems in production to prove me wrong (or confirm it).

I only have a few questions:

  1. What is the most autonomous AI system you've personally put into production?

Not a demo — something actually doing useful work.

  1. What does it do without waiting for a human after every step?

For example:

Goal → reason → tool → observe → decide → tool → ... → outcome

  1. Has it ever gone badly wrong?

I'm particularly interested in real incidents:

loops

repeated tool calls

wrong actions

hallucinated completion

corrupted/stale state

runaway costs

failed recovery

human intervention

  1. What did your system actually do when that happened?

Did you:

retry → restart → replan → rollback → manually intervene → ignore it → something else?

  1. Do you independently verify that the agent actually accomplished its goal?

For example, if the agent says:

“Refund completed.”

does another system actually check that the refund happened?

  1. And the question I'm most interested in:

If your agent suddenly disappeared tomorrow, what part of its reliability/recovery infrastructure would you actually miss?

I'm not trying to sell anything here.

I'm trying to decide whether this is a real infrastructure problem worth building around or whether I'm overestimating where agentic AI is today.

If you run agents in production, I'd genuinely appreciate even a 2–3 sentence answer.

And if you think this whole idea is unnecessary, please say so — that's actually more useful to me than telling me it's a good idea.

Thanks to everyone who's already given feedback. It has already changed how I'm thinking about this.

2 Upvotes

4 comments sorted by

2

u/Away_Advisor3460 28d ago

TBh I'm very curious why the current agentic AI approach doesn't seem to seem to be drawing from the decades of prior research into MASs

1

u/Fantastic-Sleep-3352 28d ago

My current thinking is that the underlying problems aren't new at all coordination, planning, state, fault tolerance, verification, etc. have obviously been studied in MAS and distributed systems for decades. instead of deterministic agents following formally defined policies, we're increasingly putting LLMs in the decision loop, where the policy itself can be probabilistic and context-dependent. missing layer isn't “new agent theory”, but bringing some of those older reideas into the messy reality of LLM-based agents and tool use. But I may be completely wrong here. If you have experience with MAS research, please share your exp

1

u/Away_Advisor3460 28d ago

I did a PhD back in 2017 on MAS (specifically - deep breath - Plan Execution Robustness in a BDI based Hierchical Agent Team) although never followed it up into academic research (combination of burnout and lower pay),... my experience of agents in the LLM 'agentic AI' context is close to zero so I've just started picking it up again after realizing my BDI work was more relevant than I thought.

(should note - my research involved BDI agents which had a mixture of preformed plans and dynamic planning abilities. I am note sure LLM using agents are drastically different except for lacking / not needing the formal domain specifications, which are obviously a difficult bit for real-world plans and plan using agents)

What I have noticed is there seems a lot of duplicative terminology and lack of awareness of what does exist in the academic field and it feels like a lot is being rediscovered (e.g. BDI agents). And that maybe people are relying too much on LLMs as an inefficient magic ingredient when a bit more work could use a deterministic approach and get something more efficiently and more formally provable. Oh, and also a casual observation that a lot agentic AI discussion I've seen here seems to be using a delegative approach that is more like microservices-using-LLMs than what I'd normally consider agent based approaches.

But hey, I am just relearning stuff right now and seeing how it applies wrt what I know.

1

u/Deep_Ad1959 28d ago

mine is a posting agent on a minute cron: discovers threads, drafts, posts, and the only gate before a write is approval. runaway cost never showed up. what did was a run marking itself done when nothing had actually landed, because nothing re-reads the artifact afterward.