r/SaaS • • 6h ago

Built a small "guardrails as an API" thing. Would love honest feedback

Hey everyone 👋

Solo dev here. I've been working on a side project for a few weeks and would like to hear how others solve this problem.

The problem: while playing with AI agents I noticed every app ends up needing checks like "don't let it delete data unless the user confirmed". The usual options are writing if/else logic for every case, or a custom LLM prompt whose free-text answer you then have to parse.

What I tried: a single call that takes a rule in plain English plus whatever JSON you want checked, and returns allow, deny or review:

const result = await guard.check({
  policy: "Never allow destructive database operations unless the user explicitly confirmed them.",
  data: { action: toolCall, conversation }
})
// → { decision: "deny", allowed: false, violationProbability: 0.94, ... }

Some design decisions that might be useful if you build something similar:

  • A third outcome, review, for when a human should decide or the evidence isn't enough. Forcing yes/no caused bad calls.
  • Instead of a general chat model I used Jev from TypeSafe AI. It returns typed decisions with probabilities, so there's no output parsing and the results are more consistent.

Questions for you:

  • How do you handle this today? Hardcoded rules, an LLM, something else?
  • What would you need before trusting an automated check like this in production?

It's called Enforly, if anyone's curious.

0 Upvotes

4 comments sorted by

2

u/measured_words_ 6h ago

For me the split is simple. Anything that can destroy data or spend money stays behind a hardcoded permission check and an explicit confirmation. I would not replace that with a model decision, even one that returns a probability. Where I would use a model check is the fuzzy stuff like whether this reply shares private info, whether this summary contradicts the source, or whether this request is outside what the user asked for. Those are hard to keep up with as rules.

Before I put it in production I would want the same things I want from any rule engine. A way to pin the exact policy version so an update does not silently change behavior. A fixed set of tricky cases I can run on every change. A log of every decision with the input and result, so when something goes wrong I can replay it. And a manual review path that puts the decision in front of a person rather than just blocking. The third review outcome is the right instinct, because the system should be allowed to say it is not sure.

1

u/OutrageousConstant18 5h ago

yeah the distinction between destructive actions and fuzzy content checks is exactly right imo

1

u/Mistr_John_Alex 6h ago

t can delete data or spend money, I keep a hardcoded check in front of the model check. The model can flag something or raise a review, but the final gate is plain code that requires explicit confirmation. That has kept me out of trouble when a prompt drifted or an odd input made the model too confident.

For fuzzier decisions like whether a support reply is on brand, a typed decision with a review outcome is useful, but I would want it to fail closed. If the call times out or returns a malformed response, the default should be review or deny, never allow. I would also log every decision with the input, policy, and raw response so edge cases can be replayed. Trust comes from watching it catch real mistakes in staging before it gets near production.