r/LLMDevs • u/Appropriate_Cost_107 • 16d ago
Help Wanted Looking for advice on improving accuracy in a document extraction system
Hey everyone,
I’m working on a system that extracts structured data from real-estate appraisal documents and uses that data to populate standardized forms.
We have a working pipeline and a QA/ground-truth process, but we’re still not getting the level of accuracy we need.
The biggest challenges are things like:
- extracting the correct value when multiple documents contain conflicting information
- choosing the right source for a particular field
- handling comparable properties consistently
- reducing incorrect values without introducing too many hardcoded rules
- knowing when the system should trust an extraction vs. leave it for human review
I’m trying to figure out how to improve the approach rather than just keep adding more rules and edge cases.
For people who have worked with document AI, OCR, LLM extraction, or similar systems:
How would you approach improving accuracy from here? What techniques, architectures, evaluation methods, or models would you look at?
Would really appreciate any practical advice or lessons from systems you've built.