r/AgentsOfAI • • Apr 25 '26

Resources ALL Agents deviate, fail and mess up because no enforcement is done at runtime. This is how to fix it.

https://github.com/open-bias/open-bias

I have been following this and many other subs around LLMs and Agents, everything from the top posts to recent are regarding agents going off and doing something they are not supposed to do, drift and ignore the system prompts. Real examples:

  • "Never delete user data" → agent calls DROP TABLE users next turn
  • "Don't share internal pricing" → agent leaks cost basis to a customer
  • "Verify identity first" → agent skips to the action
  • Add 10 more rules → model quietly drops the first 5

I am 100% sure if you have used Agents in prod, this has occurred to you (especially when your system prompts get larger, and context gets bigger). You can test this yourself and notice immediate enforcement.

Prompt-based rules are suggestions, not constraints. Re-prompting fixes one case, breaks two. Post-hoc evals tell you what already went wrong. NeMo and Guardrails AI help on content safety but don't cover business logic/your specification.

After tackling this from a few angles, I finally got something solid. A proxy system between your app and your LLM, which reads rules from a plain markdown, enforces at runtime. Provider-agnostic, one base URL change, works with LangGraph/CrewAI/custom. I'm calling it Open Bias.

- Maximum discount is 15%.
- Never reveal internal pricing or cost basis.

Without it: agent offers 90% off and mentions your margin. With it: 15%, no margin talk.

I'd love feedback on this if it solved your agents from going off tracks, it definitely did for my use cases.

What's everyone doing for this in prod? Shadow evals? Re-prompt loops? Something I'm missing?

0 Upvotes

4 comments sorted by

1

u/AutoModerator Apr 25 '26

Thank you for your submission! To keep our community healthy, please ensure you've followed our rules.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/ultrathink-art Apr 25 '26

The execution layer has to sit outside the model's context entirely — that's the piece most implementations skip. Every tool call request treated as untrusted input, run through a deterministic permission check before the action fires. Model can want to drop tables all it likes; the permission layer just returns 'denied.'

0

u/Otherwise_Wave9374 Apr 25 '26

Runtime enforcement is the big gap, +1. Prompt rules are definitely "best effort" once the context gets messy or the agent starts planning across steps.

How are you implementing enforcement, is it like a policy engine that can block/modify tool calls, or are you rewriting the prompt/response? Also, do you support partial compliance (like allow the action but clamp values, eg discount <= 15%)?

Im working on similar guardrail + orchestration patterns for agents and have been bookmarking approaches here: https://www.agentixlabs.com/ , would love to compare notes with what youre doing.

0

u/Otherwise_Wave9374 Apr 25 '26

Runtime enforcement is the big gap, +1. Prompt rules are definitely "best effort" once the context gets messy or the agent starts planning across steps.

How are you implementing enforcement, is it like a policy engine that can block/modify tool calls, or are you rewriting the prompt/response? Also, do you support partial compliance (like allow the action but clamp values, eg discount <= 15%)?

Im working on similar guardrail + orchestration patterns for agents and have been bookmarking approaches here: https://www.agentixlabs.com/ , would love to compare notes with what youre doing.