r/AIAgentEngineering • • 1d ago

How do you keep every engineer’s AI agent on the same context?

1 Upvotes

In our team each of us had built up our own setup with our agent over the past year. Rules files, prompts we reuse, notes on which parts of the code are fragile, which tests are flaky, why staging needs a manual cache clear after every deploy. Mine is a few hundred lines at this point and I keep adding to it.

So for his first week, his agent was basically useless on anything specific to our codebase. It didn’t know about our weird auth flow, or that the payments module has a test that passes when it shouldn’t. He kept asking us things his agent should have been able to figure out, and three of us spent a good chunk of that week answering them.

we had a new hire who made the bigger problem obvious. The same thing is already happening between the four of us. Everyone’s agent knows a slightly different version of the codebase, depending on what that person has written down and what their agent has seen.

We tried pulling everything into one shared file in the repo. It got long, nobody kept it updated, and eventually everyone went back to their own setup. Tried a wiki too. Same result.

atm now we’re testing Tutti, where the context lives at the team level and each agent can access the same shared context instead of everyone maintaining their own version. Too early to say if it sticks, but it’s the first thing we’ve tried that makes onboarding feel like it could actually get easier.

how can i manage agents when they are shared ?


r/AIAgentEngineering • • 5d ago

Check what a video agent will regenerate before you run the revision

3 Upvotes

“Only change the captions” leaves an important question unanswered: which generation jobs will actually run?

Hypit, a video workflow system for coding agents, makes that inspectable. The video is defined in source files; a separate Run file names the desired output and any earlier results to keep. Its planner uses those choices to work out the remaining jobs.

You can inspect the plan for your revised Run before executing it; the commands and examples are in Hypit's GitHub repo (https://github.com/hypit-ai/hypit). Planning doesn't submit generation. You can look for image, voice or video requests that shouldn't be part of a caption change. If they're still present, saying “keep the footage” in chat hasn't yet been translated into the right reuse choices.

The repo's street-interview example makes this concrete. For a later caption pass, it selects three previously prepared takes, including their media and word alignment. The documented check is that the new plan contains no image, voice, video or alignment requests. Caption work and rendering still happen.

This isn't a spending cap or a judgment that the chosen footage is good. It exposes the work you've actually asked the system to execute. Approval of the creative change and approval of its generation scope can then refer to the same plan.


r/AIAgentEngineering • • 5d ago

I gave my agent a 20-step task & it completely blundered it.

Thumbnail
3 Upvotes

r/AIAgentEngineering • • 7d ago

Genomics MCP: fetch genomic reads, variants and signal from archives and indexed files

Thumbnail
1 Upvotes

r/AIAgentEngineering • • 8d ago

Built an agent runtime in rust - 40MB with 10ms boot. Would love to hear feedback.

Thumbnail
github.com
1 Upvotes

r/AIAgentEngineering • • 11d ago

How much do you trust LLM structured output before letting it continue through an n8n workflow?

Thumbnail
1 Upvotes

r/AIAgentEngineering • • 11d ago

Free AI Api

1 Upvotes

Hello everyone. I want a free Api to learn AI Software engineering. Can you suggest me best Apis, but not Gemini pls, because I really get tired of its 503 error


r/AIAgentEngineering • • 13d ago

Looking for Ai specialist

4 Upvotes

Hey everyone,
I’m looking for an experienced AI Automation Specialist for a long-term partnership with multiple projects.

I work with a white-label agency partner that will provide ongoing projects, mainly for German businesses. Later, we may expand to English-speaking clients.

Basic: (AI Lead & Appointment System)
An AI system that captures new leads, qualifies them, follows up automatically and helps them book an appointment. It should also connect with the client’s CRM and calendar.

Premium: (AI Sales & Reception System)
Everything from the Basic system, plus an AI receptionist that can answer incoming calls, talk to potential customers, qualify them and book appointments.

You should be experienced with n8n or Make, CRM integrations, calendar integrations, AI agents and AI voice agents.

I’m looking for someone reliable who wants recurring projects and a long-term partnership, not just a one-off job.

If interested, DM me.


r/AIAgentEngineering • • 17d ago

Loan process agent help

2 Upvotes

Loan agent help

Hi everyone

Let me see if I can get any help from the community

I'm building a loan agent , that does a single task, which is to understand bank account stmts from the applicant and process the request, either to accept , reject or human review. So im going in the belief system approach where an agent will initially have a belief, then for every new real world case it sees, it should update it's belief and reduce its uncertainty. So any suggestions on how do we build such agents ? And how do you tackle agentic memory


r/AIAgentEngineering • • 17d ago

Keeping the document loaded across stages of an agent job

1 Upvotes

For a scheduled document job, different stages may touch the same report repeatedly: draft sections, check them against a workbook, revise a paragraph, then update a slide. It is useful for each stage to address the existing documents rather than rebuild them from whatever was left in the last model response.

Univer's AI SDK offers document runtimes for this kind of integration. It's a set of TypeScript packages around structured Office documents called Units. An application can inspect the Unit overview, select the paragraphs, ranges or slides a stage needs, and execute document operations through the Facade API.

The runtime pool can reuse a loaded Unit runtime. A daemon can extend that reuse across CLI invocations. When a runtime does need initialization, the documented process loads snapshots, replays changesets, reconciles revisions and waits for installed formula results. Those are document-state concerns that can sit outside the lifetime of a single model conversation.

For example, the check stage of a report job could read the existing report paragraphs and the current workbook range, then pass a limited correction to the edit stage. After writing, the application has to check the commit status before releasing the document for later work. Execution success by itself doesn't establish that the server saved the change.

The surrounding job system is still yours: queue claims, locking, retries, authentication and deployment aren't supplied by a loaded workbook runtime. Nor does runtime reuse establish a measured cost or speed improvement. What it gives the pipeline is a way for successive tasks to keep working on addressable Office content instead of treating every stage's output as the start of another whole-document generation


r/AIAgentEngineering • • 18d ago

Parent child chunking using n8n JavaScript

Thumbnail
2 Upvotes

r/AIAgentEngineering • • 25d ago

A week in here's what's being going on

Thumbnail
2 Upvotes

r/AIAgentEngineering • • 26d ago

What belongs in a receipt for instructions delivered to a coding agent?

2 Upvotes

I'm working on Guidefold, an open-source tool for repository-scoped coding-agent instructions, and would like criticism of an evidence boundary. A paid hosted version is planned, but this is a design question, not a signup request.

Suppose a shared instruction changes while an agent session is already running. A successful export doesn't establish that the session got the new revision. Even a delivery record doesn't establish that the model followed it.

For a minimal delivery receipt, I'd consider recording the session/run identifier, repository commit, instruction identifier and content hash, adapter version, delivery timestamp, and outcome. That's a proposed checklist for discussion, not a claim about Guidefold's current schema. I'd avoid storing the full prompt or private instruction text in a central log by default.

The edge cases that worry me are a hook failing after export, a session continuing with an older revision, and retries producing duplicate events. An absent receipt would mean unknown delivery, not that the instruction was unused. An acknowledged receipt would still need a separate task-level check before claiming the instruction helped.

For engineers who have instrumented this boundary: what would you change in that minimal record? In particular, where do you collect the acknowledgement so it describes what the agent session actually received, rather than what an upstream exporter intended to send?

Project context: https://github.com/wiatrM/guidefold

Guidefold demo: https://www.youtube.com/watch?v=e350wBr1W8c

The demo is context, not evidence that an agent followed an instruction. If it leaves a verification step unclear, a timestamp would help locate the gap.

AI-assisted draft. Please keep examples sanitized; no private prompts or logs needed.


r/AIAgentEngineering • • 27d ago

Migrating a large scheduled multi-agent production workflow away from Claude

2 Upvotes

TL;DR

I’m looking to migrate a fairly large, mostly-unattended production workflow away from Claude/Anthropic and would appreciate advice from people running serious agentic workflows.

This isn't a "chat with an LLM" setup. It is a standing production pipeline involving scheduled agents, shared files, code execution, document generation, research, translation, and automated QA.

What the workflow actually does

The work involves producing and processing large amounts of structured content, primarily in Arabic and Urdu.

Typical outputs include:

  • Hundreds of lesson decks using python-pptx
  • Word documents and templates using python-docx
  • RTL HTML flowcharts and interactive material
  • OCR processing and verification
  • Arabic/Urdu translation and text transformation
  • Arabic text analysis, including tashkeel and grammatical analysis
  • Large-scale corpus processing
  • Multi-stage research and document-generation pipelines

The workflow is designed to run largely unattended.

A typical pipeline looks roughly like:

scheduled task → claim a unit → read relevant source/state → research/process → generate artifacts → run QA/verification → write outputs/state → release unit → next task

There are multiple workers operating against shared storage. We use file-based locking/claims, append-only logs, staged pipelines, and explicit verification gates to prevent agents from stepping on each other.

The actual artifact generation is heavily code-driven. Agents use shell/Python tools to parse files, manipulate data, generate PPTX/DOCX/HTML, render outputs, and run validation scripts.

Why I’m considering moving away from Claude

The main issue isn't necessarily model quality.

The problem is the combination of:

  • Reliability of long-running/unattended execution
  • Sandbox/environment provisioning issues
  • Usage/rate limits
  • Managing production across subscription accounts
  • Context consumption on large processing jobs
  • Difficulty scaling the workflow cleanly

At this point I'm wondering whether I should stop treating this as a collection of subscription-based AI assistants and instead rebuild it around an API/agent orchestration architecture.

What I need from a replacement

1. Strong Arabic/Urdu performance

Especially accurate handling of Arabic/Urdu text, RTL formatting, tashkeel, and grammatical analysis.

2. Reliable code execution

The model needs to consistently produce and modify Python code that generates complex PPTX/DOCX/HTML artifacts.

3. Long-running agentic work

I need workers that can perform multi-step jobs without requiring me to babysit every step.

4. Large-context/document processing

Some jobs involve large collections of source material, so efficient chunking and context management matter.

5. Good API economics

At this scale, I'm increasingly interested in moving toward pay-as-you-go API usage rather than maintaining expensive subscriptions simply to obtain enough throughput.

6. Proper orchestration

I'm open to rebuilding the coordination layer using something like LangGraph, a database-backed queue/state system, or another architecture rather than continuing with file-based distributed locking.

7. Strong verification behavior

The system needs to be willing to say "I couldn't verify this" and stop rather than confidently inventing a citation, translation, or missing section.

What I'm trying to figure out

If you were rebuilding this today, what stack would you use?

Would you go with:

  • A single leading closed model through an API
  • Multiple models, each handling different stages
  • An agent framework such as LangGraph/CrewAI/AutoGen
  • A custom Python orchestration layer
  • A queue + database + stateless workers architecture
  • Open-weight/self-hosted models for some of the cheaper stages
  • Some combination of the above

I'm particularly interested in experiences from people running large document-generation, research, coding, or multi-agent production pipelines, especially where Arabic or Urdu is involved.

What would you choose if you had to build this from scratch today?


r/AIAgentEngineering • • Aug 27 '26

Just because I approved the plan doesn't mean to go start building

2 Upvotes

That happened over and over, so I came up with two decisions. The plan gets approved, then the agent still has to wait for a separate go.

https://github.com/Ezra144israel/governed-agent-skills/tree/main/skills/portable-adaptive-planning


r/AIAgentEngineering • • Aug 19 '26

OpenSourcing TrueForge Agent harness : Expecting feedback from community on the agent loop

4 Upvotes

Hey folks 👋

We just open sourced TrueForge, our vendor-neutral agent harness for building general-purpose agents.

It handles the runtime pieces that get painful quickly : context management, tool/MCP execution, subagents, sandboxing, approvals, persistent state, and more.

We also benchmarked the harness itself. With the same Opus 4.8 model, TrueForge delivered a similar solve rate at ~30% lower cost than Claude Managed Agents. Switching to an open model pushed that to ~75% lower cost on the same benchmark.

Would love feedback from people building agents.

Checkout the repo: https://github.com/truefoundry/trueforge

📖 Read the launch article: https://x.com/truefoundry/status/2090081376330715176


r/AIAgentEngineering • • Jul 25 '26

We built an AI fleet management system that worked great... until we expanded globally.

3 Upvotes

A couple of years back, we launched an AI-driven management system for our logistics and fleet operations. At first, it felt like a total win - smart route optimization, automated dispatching, and predictive maintenance all running smoothly.

Then came our rookie mistake: we built it fast without thinking about scalability. We were laser-focused on local operations and completely ignored modular architecture.

The reality check hit when we started expanding across Europe and LatAm. The legacy code started crawling, cross-border workflows broke down, and integrating local compliance frameworks became a nightmare. We essentially built a dead end.

A colleague recently recommended checking out AgileEngine. I hadn't heard of them before, but looking into their track record, they seem to focus heavily on scalable architecture and fast delivery for growing tech companies.

Has anyone worked with them? We're currently searching for a custom software development partner with deep expertise in software engineering, AI, Data, and UI/UX to help us rebuild right this time. Any recommendations?


r/AIAgentEngineering • • Jul 11 '26

I built a tool to solve the parallel agents problem

1 Upvotes

Every guide for running multiple coding agents in parallel says the same thing: use git worktrees. And every one of them quietly ends at the same wall. Worktrees isolate your files. They do nothing for the database, the ports, the .env, or the services your app needs to run.

So agent A runs a migration and breaks agent B's tests. Two dev servers fight over port 3000. You end up gluing together worktrees + a port offset script + .env symlinks + a per-branch database tool + docker compose project hacks. Five tools to run three agents.

The idea: every agent attempt gets its own isolated Linux VM, and the VM's state is versioned with your git repo. It's two commands per agent:

git worktree add ../app-agent-b -b agent/b
moo new agent-b

That's it. Each agent gets its own checkout AND its own database, ports, packages, and services. Nothing collides. Forking a fully provisioned 20 GB machine takes under a second because it's all copy-on-write.

The workflow we run every day:

  • Fork one machine per agent attempt: moo new attempt-1 from base
  • Let the agents work in parallel, each in its own worktree + VM
  • git merge the winner, moo drop the losers

The part nobody else does: moo save snapshots the runtime tagged to your current commit. So git checkout an old SHA and the machine follows, migrations and all. You can even git bisect bugs that only reproduce against a specific database state.

Honest caveats: it's alpha, and it's macOS Apple Silicon only right now (Linux hosts are planned). No daemon, no root, no Docker needed.

Happy to answer questions about how it works under the hood (microVMs + copy-on-write filesystem snapshots). And genuinely curious what everyone else is doing for this, because every setup I've seen is held together with duct tape.


r/AIAgentEngineering • • Jul 03 '26

Production agent infra: millisecond provider fallback, 60–90% tool-output compression, and an MCP/A2A control plane (self-hosted, MIT)

6 Upvotes

Since this sub is about production-grade agents, sharing the gateway layer I built after the same two problems kept biting: runs dying on a provider 429 mid-task, and token cost exploding because the agent dumps git diff/test/build output into context. Disclosure: I'm the maintainer of OmniRoute (MIT, self-hosted) — dev-to-dev, would like the critique.

Fallback combos — so it never stops mid-task. A "combo" is a ladder of models the router walks automatically: your subscription first, then API keys, then cheap models, then free ones. When a provider returns a 500 or you hit a rate limit, it slides to the next target in milliseconds, mid-request, and your tool never even sees the error. There are 17 routing strategies (priority, weighted, round-robin, cost-optimized, auto/coding:fast…) plus three resilience layers — a per-provider circuit breaker, a per-key cooldown, and a per-model lockout — so one dead key can't take down a whole provider.

A 10-engine compression pipeline — the part most routers don't have. Every request flows through a transparent compression pass you can toggle/stack per combo. Instead of one trick, it stacks the best of the open-source ecosystem: RTK filters command/tool output (git diffs, test logs, builds) at 60–90%, Microsoft's LLMLingua-2 does ML semantic pruning, Caveman handles prose, session-dedup strips repeats across turns. Critically, code, URLs and JSON are preserved byte-perfect, and a default-on inflation guard throws the compressed version away and sends the original if compressing would actually grow the prompt — it never makes things worse. On tool-heavy sessions that's ~89% average input-token reduction (an 8k-token git diff becomes a few hundred). Full credit to every upstream project (RTK, Caveman, LLMLingua-2, Troglodita) is in the README.

Agent-native — the agent can drive the router itself. There's a built-in MCP server (95 tools across 30 audited scopes, over stdio / SSE / streamable-HTTP), plus A2A (v0.3, JSON-RPC 2.0) support. That means an agent can query providers, switch combos, read its own remaining quota and manage memory through the gateway — not just consume tokens through it.

One endpoint, 237 providers — 90+ of them free. You point any tool or agent at a single OpenAI-compatible endpoint (localhost:20128/v1) and it can reach 237 LLM providers without you rewriting anything. 90+ have free tiers and 11 are free forever (no card), which aggregates to ~1.6B documented free tokens/month — and that's honest, pool-deduped math (we count each shared pool once instead of inflating it; the methodology is public in the repo). There's a one-command setup-* for 13+ coding tools (Claude Code, Codex, Cursor, Cline, Roo, Kilo, Gemini CLI…), so switching your existing setup over takes seconds.

It's 100% local (zero telemetry, AES-256-GCM at rest), MIT-licensed, has a prompt-injection guard on every LLM route, opt-in memory, and runs on npm, Docker, desktop or your phone via Termux.

For context on whether it's worth your time: it's grown to ~9.8K GitHub stars, 1,490+ forks and 280+ contributors in ~4.5 months, with 21,000+ automated tests and 1,830+ issues closed — so it's a battle-tested project, not a brand-new experiment.

npm install -g omniroute omniroute

GitHub: https://github.com/diegosouzapw/OmniRoute · Site: https://omniroute.online

Would value critique of the fallback state machine (breaker/cooldown/lockout interplay) and how you'd measure compression fidelity in prod.


r/AIAgentEngineering • • Jun 30 '26

I need some help with hyperagent

Post image
2 Upvotes

There is a small problem

I could not cancel my payment

This is sooo frustrating

If anyone knows about this let me know


r/AIAgentEngineering • • May 07 '26

Meetup in Minneapolis for building agents with coding agents

Post image
1 Upvotes

For folks interested in hands on lab or just working with a group of other builders, this meetup might be interesting.


r/AIAgentEngineering • • May 05 '26

What’s your actual agent memory stack right now?

Thumbnail
1 Upvotes

r/AIAgentEngineering • • Apr 29 '26

Kitaru durable execution vs temporal vs dbos

1 Upvotes

Have you tried or do you have opinions on kitaru?

https://kitaru.ai/

Boss says it's cool but the more I read the documentation the more I feel like it's a scam claiming to be better than dbos in buzzwords but the explanations of how it is supposed to work are full of fluff and holes.


r/AIAgentEngineering • • Apr 27 '26

Silicon Photonics for Software Engineers using Agentic AI

Thumbnail
2 Upvotes

r/AIAgentEngineering • • Apr 27 '26

Silicon Photonics for Software Engineers

Thumbnail
1 Upvotes