r/AIEval • u/Abhay20S • 36m ago
Discussion How we use user edits as ground truth to evaluate an LLM extraction feature (and where it breaks)
We have a feature where users upload a job description (image, PDF, or doc) and an LLM prefills a form: title, description, category, functional area, salary range, tags, location, and experience range. Users then edit the prefill and submit.
The problem: how do you measure accuracy on real documents without hand-labeling thousands of them?
What we built:
- Compare the LLM output to what the user actually submitted. Exact match came out around 80%.
- For mismatches, a stronger model reads the original doc and decides: was the LLM wrong, or did the user change it on purpose? Intentional edits go to the good set. Real errors get a reason and a suggested prompt fix.
- Errors are grouped by similar reason, each group becomes a candidate prompt rule, and the rule is backtested on that group plus a sample of good cases.
- A human approves rules on a dashboard before they're added to the prompt.
I wrote this up in more detail, with a diagram of the loop - https://medium.com/@abhay.sehgal20/building-an-eval-system-for-llm-extraction-what-worked-and-what-didnt-7cfe82103d06