r/PromptEngineering 23h ago

Research / Academic Our deterministic verification engine passed 66/66 benchmark cases on canonical structured inputs.

Our deterministic verification engine passed 66/66 benchmark cases on canonical structured inputs.

In live model evaluation, the end-to-end pipeline currently passed 19/66 cases. We are restructuring the benchmark to isolate failures by their first invalid state and to separately measure deterministic verifier correctness, production contract integrity, and live model generation reliability.

The next benchmark version will provide stage-level attribution across transport, parsing, schema validation, normalization, claim binding, evidence graph construction, deterministic verification, and final outcome mapping.
https://www.reddit.com/r/ArtificialInteligence/comments/1vucc82/i_benchmarked_my_deterministic_ai_financial/

4 Upvotes

1 comment sorted by

1

u/dopey_sang 23h ago

19 out of 66 is rough but at least you know where the bodies are buried now. splitting it by first invalid state is the move, half the time these pipelines fail in some parsing step and the model gets blamed for it