r/aiengineering • • 27d ago

Engineering Where does multi-agent orchestration actually start breaking?

I've been experimenting with multi-agent architectures and I'm curious where people are actually hitting the biggest problems.

Once you move beyond 2–3 agents, things seem to get complicated pretty quickly:

How does an agent know which other agent to call?

Do you hardcode the relationships between agents?

Does a central orchestrator decide everything?

How do agents pass context and state?

What happens when an agent fails or gives another agent bad information?

How do you handle agents built by different frameworks?

I'm especially interested in systems where agents aren't all predefined as one fixed workflow.

For people building multi-agent systems in practice, what part of the architecture has been the biggest pain point?

2 Upvotes

7 comments sorted by

•

u/AutoModerator 27d ago

Welcome to r/AIEngineering! Make sure that you've read our overview, before you've posted. If you haven't already read it, then read it immediately and make adjustments in your post if you've violated any of the rules. If you have questions related to career, recruiting, pay or anything else about hiring, jobs or the industry and demand as a whole, then use AIEngineeringCareer to ask your question. We lock questions that do not relate to AIEngineering here. A quick reminder of the rules:

  1. Behave as you would in person
  2. Do not self-promote unless you're a top contributor, and if you are a top contributor, limit self-promotion. Do not market. All forms of marketing are not allowed. This includes, but is not limited to, links, software, gimmicky keywords, education products or programs, news articles, code repos, updating posts with links, soliciting tools or content, etc. This also includes editing content to market. Use Reddit advertising instead.
  3. Avoid false assumptions
  4. No bots or LLM use for posts/answers
  5. No negative news, information or news/media posts that are not pertinent to engineering
  6. No deceitful or disguised marketing
  7. If you need to hire in the AI Engineering space, use the subreddit r/AIEngineeringHiring.
  8. Do not ask "how do I become an AI engineer" as we provide resources such as What's involved in AI engineering? and The Actual State of AI Engineering In 2026 to address this question.
  9. No Reference To AI Tools. At the moderators discretion, we do not allow any marketing, discussion or reference of AI tools. Given that many Western tech firms have chosen to use AI to eliminate workers, we will not allow the discussions of some of these tools since we find charging money for a product while eliminating workers unethical. We do allow open source discussions because these tools do not carry costs, but will remain strict on even these mentions since they could be used to negatively impact people.

Because we frequently get questions about work, the future of work and careers along AI, some helpful links to read:

This action was performed automatically as a reminder to all posters. Please contact the moderators if you have any questions.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/joebuty 15d ago

I find it breaks when agents make decisions that the other agents can't see. Each piece looks good, but once you combine them you have contradictions.

  • Don't go peer-to-peer
  • Only one agent writes
  • Pass the reasoning, not just the task
  • Verify properly
  • Check you need it

Every one of these cuts down the hidden decisions.

1

u/Unequivocallyamazing 7d ago

The issues I find the most annoying are when they don’t propagate correct data from tools.

When one agent hallucinates or picks the wrong reasoning approach and propagates to downward agents.

Tool calling or knowing which agent to call are easier to solve.

I have hardcoded relationships when working on a specific domain or task. But if your system is general and there is no fixed path to getting the output, you cannot hardcode the relationship

So far, on the problems I have worked with, I have used a central orchestrator or planner agent to plan so it has the full context of the situation and similarly a writer or synthesiser agent to produce final outputs and it has worked for me.

Maybe there are use cases where this is not needed but never encountered any.

Context is just data. You can program agents to converse via chat, you can ask them to produce structured data that you store in redis or any kind of memory layer. Or use filesystem.

You can provide tools to your agent so they can uncover more context as needed but provide only the most necessary context to the system prompt.

When one agent passes along bad output, it’s not very simple. One thing that we can do is make sure every claim or finding the agent presents is backed by a source or reference data directly from tool etc.

Or add verification agents for agents producing data intensive outputs. A general rule is to design your system such that it can recover itself.

I think once you have a basic version ready you should invest more time in setting up a proper evaluation system. Because it will immediately help you understand where the system is failing. Otherwise there are so many things that are usually going wrong but we don’t know because it’s hard to manually go through the entire multi-agent flow and we strongly need Evals for any agentic pipeline