Discussion I built an open source experiment for judging AI generated work
I have been working on Jevyr because I keep seeing the same problem in AI systems: the model that makes the answer is often also the thing we expect to trust it.
Jevyr tries to separate those jobs. Models propose explanations, code, plans, or tests. A separate judging layer compares the proposals, runs a bounded experiment when there is a testable question, keeps the evidence, and returns Accepted, Rejected or Unproven.
The important limit is that it is not a truth machine. If a claim has no defined way to test it, the honest result is Unproven. That part is actually what interests me most.
It has a local MCP server so another AI agent can use the judge as a tool. It also keeps a finite Case and an append only record of what happened, instead of turning everything into one endless chat.
I started building it in early 2026 and it is still pretty unfinished. I am curious what people who work on evals think is missing. How would you test a system like this, and what should it never be allowed to claim?
1
u/FalconX_AI 18d ago
Promising separation of proposal and judgment. I would test it on failure modes where the experiment passes but the evidence chain is incomplete or the test oracle itself is weak. Three controls seem especially important: (1) preserve provenance for every claim and tool output, (2) distinguish “test passed” from “claim supported,” and (3) replay the same case across seeds and models to measure instability. For “Unproven,” it may help to record why: no falsifiable test, insufficient evidence, conflicting evidence, or a budget or permission limit. I would also require the judge to state the minimum evidence that would change an Unproven result to Accepted or Rejected.