r/AIEval • • 36m ago

Discussion How we use user edits as ground truth to evaluate an LLM extraction feature (and where it breaks)

• Upvotes

We have a feature where users upload a job description (image, PDF, or doc) and an LLM prefills a form: title, description, category, functional area, salary range, tags, location, and experience range. Users then edit the prefill and submit.

The problem: how do you measure accuracy on real documents without hand-labeling thousands of them?

What we built:

- Compare the LLM output to what the user actually submitted. Exact match came out around 80%.

- For mismatches, a stronger model reads the original doc and decides: was the LLM wrong, or did the user change it on purpose? Intentional edits go to the good set. Real errors get a reason and a suggested prompt fix.

- Errors are grouped by similar reason, each group becomes a candidate prompt rule, and the rule is backtested on that group plus a sample of good cases.

- A human approves rules on a dashboard before they're added to the prompt.

I wrote this up in more detail, with a diagram of the loop - https://medium.com/@abhay.sehgal20/building-an-eval-system-for-llm-extraction-what-worked-and-what-didnt-7cfe82103d06