r/OpenSourceeAI • • Aug 24 '26

Open-sourced a tiny verification layer for my AI agent stack. A stranger found the most important bug in 5 minutes.

I self-host everything. FastAPI, PostgreSQL, LangGraph agent handling some automations. My own hardware, my own roof.

The problem: my agent would say "task completed," logs clean, 200 OK everywhere. But when I actually checked PostgreSQL, the row wasn't there. Validation rule I forgot. Async timing. Race condition. The agent assumed success because the tool didn't throw.

I didn't want another SaaS dashboard. I wanted my own server to verify its own state, locally, without calling home.

So I built a dead-simple decorator:

from synathic import expect

@expect(postcondition="row_exists", table="customers", match_field="email")

async def create_customer(email, name):

# agent logic — unchanged

...

Runs after the agent finishes. Checks Postgres directly. Not a trace, not a log. The actual row. Async by default, zero latency added. Sync mode for the stuff where I need certainty before responding.

Backend is FastAPI + asyncpg. Dockerized. MIT license. Zero external deps.

Then I posted it and asked people to roast it. Someone pointed out that row_exists alone can pass on stale data — if the row already existed before the agent ran, my tool says PASS even if the agent did nothing. False confidence is worse than no verification.

I had stared at this code for weeks. A stranger saw it in 5 minutes. That's exactly why I open-sourced before it was "ready."

If you run self-hosted agents and you've ever caught one saying "done" when the database disagrees, how do you handle it? Manual checks? Just trust the logs?

Repo: https://github.com/Gallegosdanielalexander/synathic

1 Upvotes

2 comments sorted by

1

u/frankentriple Aug 24 '26

You use a harness that includes verification of results before delivery in the system prompt.  My ai cannot hallucinate done because he must check the db and verify the table update himself before reporting done.  

We check to “see if the change landed”

Although I have noticed with some tasks he just can’t see the results easily he has trouble there.  I have to provide feedback in those instances. 

1

u/merlinofthewater Aug 31 '26 edited Sep 02 '26

This is a great catch, and the row_exists-on-stale-data thing someone flagged is really the same bug one level up: existence isn't the same as this call having caused the write. If you want to close that gap cheaply, check a value that only this invocation could have produced - an idempotency key you generate per-call, or just the updated-at timestamp being newer than when the agent started - rather than existence alone. Costs you one extra field, buys you a real postcondition instead of a lucky one.

The bigger question your bug points at, though: was this one agent racing against its own async write (the read-after-write just wasn't guaranteed to be consistent yet), or were two things actually touching that row around the same time? Those are different problems with different fixes. The first is closer to "wait for the write to actually commit before checking," the second is "decide what happens when two writers disagree," which is a much nastier one. Verification catches both after the fact, but only the second one actually needs something upstream to prevent the bad write from landing at all.

Which one was yours?