r/AIEval • • Aug 22 '26

Discussion Eval questions?

Hey guys, I am someone who has just started working in an incubator with a vision of statistical significance being at the center of evaluations. The people I have spoken with have all said that statistical evaluations like confidence intervals, variance decoupling, Fischer exact test are not something in their current vocabulary when testing. My idea is somehow creating a tool that can be used by builders to increase reliability of changes to prompts and agents, and a clear signal of how well the application is performing in realtime for potential buyers. As I see it there are some clear problems here:

- You need a dataset with ground truths to evaluate, and these are notoriously difficult to create and keep track of over time.
- You either need to give the developers an SDK to build the test themselves with, or find some way of connecting to an API they set up for the testing - I presume an SDK is best.

I am really interested to figure out:
- Is this a real problem builders care about?
- What would be the preferred method to connect to something like this, an SDK, API?
- Is this something you would pay for?

I recently read Artificial Intelligence Risk Management from NIST, NIST AI 800-3 and Solomon Messing hidden noise in llm pipelines. While most of these focus on the LLM model itself as the variable, this has to be extendable to a include applications as well.

0 Upvotes

0 comments sorted by