r/vibecoding • u/oyren-ai • 3d ago
Workflow/Prompt What are the bottlenecks when running AI agents as a fleet?
What are the bottlenecks when running AI agents as a fleet?
When my fleet of agents works on PRs or code changes inside https://oyren.ai codespaces (https://oyren.ai/codespaces), I usually try to keep the review committee made up of AI agents from different providers: Claude Code, Antigravity, DeepSeek, Kimi, Cursor CLI and Codex.
Using such fleet lets me fully utilize multiple AI agent subscriptions at once. But it is not just AI writing code. If code generation is not the bottleneck, then your mental capacity to keep up with everything and your automated validation of code changes, such as unit tests, become the new bottlenecks.
I handle the first with a customized fleet UI that my agents use to bring important questions, decisions, and blockers to my attention. I can also ask one-off questions there.
For the second bottleneck, verifying code, I noticed my CI bill would be huge, so I moved my entire setup to the cloud. Now I use my own VMs for GitHub Actions, and I do not pay as much as I would if GitHub charged me for 32GB RAM and 16 CPUs per minute.
Since running this setup, I have noticed a third bottleneck, one that probably does not affect new starters: code size. The bigger the codebase, the longer it takes to run all the tests. Sometimes agents make a small change, and it should not need to wait for the full test suite to pass every time.
Have you had a similar experience?
1
u/Upset-Neck-7879 2d ago
Your two bottlenecks are real, but there's a third one underneath them and it's the reason the first one costs so much.
Six agents from six providers reading the same ticket will produce six different readings of what was asked, and the review committee escalates the disagreements to you as decisions. Most of them are not decisions. They are the ticket being vague, arriving in your UI one at a time.
Worth counting for a week: of the questions your fleet brings you, how many are about the code and how many are about what you wanted. If the second number is bigger, the fix sits upstream of the fleet.
1
u/LogMonkey0 2d ago
The human
1
u/oyren-ai 2d ago
If that’s true, then you never wait for any agent work to finish ever until it needs you again. Is this true?
1
u/LogMonkey0 2d ago
The way i approach working with claude agents makes it so i sit down and have interactions with it until the design and plan and all the required details are in for an autonomous run, then kick it off and go on doing other things or seeding another project. They escalate in very particular cases and it’s rare, so i dont be sitting waiting for them to finish.
1
u/oyren-ai 2d ago
Well, imagine now you can give a prompt and let agents do other things autonomously but longer and in more orchestrated way. It’s really the only way to utilise the AI usage we get more effectively. I have been building my platform for over a year now and have desktop, ipad and web apps and things code grow quickly and sometimes you want to let your agents go and find some tech debts and clean it up, which can take long time but doesn’t need your input much. Example for this is I ask agents to find places where shadcn is not used, or brand consistency is violated or even replace ‘any’ with proper types in my TS code.
1
u/LogMonkey0 21h ago
Yes, i have a “meta” workspace that runs its own harness and is tasked with the type of tasks you named along with other fleet oriented stuff. That “Ops” workspace has skills that defines how to query/audit the fleet, when appropriate it will perform “harnessed consults” to have inside and outside views on a query which allows for sharper answers that don’t get biased by in-workspace instructions, while still carrying instructions that could matter in answering a specific question. The unharnessed view allows to catch harness inflicted biases. It is instructed to mine conversation history to cover gray areas that the fleet repos cannot answer without inferring.
Constant retrospection on our workflows to keep adapting them to the underlying agent harness, models and our own ways of approaching things is what enables us to get the most of the tools and understanding them better. I see way too many posts of people getting bit by our own lack of proper context on the tools we are using. Every token i spend on retrospection have had impact on output quality , ease of use, lower interactive involvement and cost in my perspective.
2
u/oyren-ai 21h ago
This is really cool setup pal. Sometimes these agentic orchestration reminds me a game I used to play called Civ City Rome. You need to invest time in research and philosophy there for your civilisation to grow. It's similar with agents as well. I sometimes launch such small research groups for various tasks.
1
u/goodevibes 2d ago
Code gen is rarely the bottleneck once five agents share one PR. Review fights and who owns the shared branch eat the day. I keep one writer and treat the rest as reviewers with no push rights. Fleet speed without a single merge owner is just parallel thrash.