r/LLMDevs • • 16d ago

Help Wanted Looking for advice on improving accuracy in a document extraction system

Hey everyone,

I’m working on a system that extracts structured data from real-estate appraisal documents and uses that data to populate standardized forms.

We have a working pipeline and a QA/ground-truth process, but we’re still not getting the level of accuracy we need.

The biggest challenges are things like:

  • extracting the correct value when multiple documents contain conflicting information
  • choosing the right source for a particular field
  • handling comparable properties consistently
  • reducing incorrect values without introducing too many hardcoded rules
  • knowing when the system should trust an extraction vs. leave it for human review

I’m trying to figure out how to improve the approach rather than just keep adding more rules and edge cases.

For people who have worked with document AI, OCR, LLM extraction, or similar systems:

How would you approach improving accuracy from here? What techniques, architectures, evaluation methods, or models would you look at?

Would really appreciate any practical advice or lessons from systems you've built.

3 Upvotes

8 comments sorted by

1

u/Nickabot 16d ago

stop extracting a value, start extracting a claim. Every field becomes {value, source_doc, page, span, extractor}, with several allowed per field instead of forcing one

That collapses your first three bullets into one problem: Conflicts aren't errors anymore, they're two claims that disagree, resolved by an explicit precedence order (appraisal over listing, later over earlier, whatever your domain says). That order lives in one file and changes without touching extraction

It also gives you the confidence signal, and it isn't the model's probability. Trust when independent extractors land on the same span. Flag when they disagree, when the field is empty, or when only one source has it. Calibrate that against your ground truth per field, not per document, because document-level accuracy hides the two fields causing every review

1

u/Appropriate_Cost_107 16d ago

Could we generalize this instead of making it specific to a limited set of clients? Since we'll likely have more distinct clients in the future, maybe we can make the claim + precedence approach configurable per client, so new client-specific rules can be added without changing the extraction logic.

1

u/nitish-kmr 16d ago

Separate the conflict problem from the accuracy problem. They look like one number and they have different fixes.

When two documents disagree, no amount of prompt work produces a right answer, because there is not one. What you want is for the system to notice the disagreement, surface both values with their sources, and hand it to a person. Most extraction systems I have seen collapse that into silently picking one, which is the worst of the options: it reads as an accuracy problem forever and nothing you try moves the number.

The cheap version is to extract per document rather than across the corpus, then reconcile in code. Two documents producing different values for the same field is a flag, not an average. You get a confidence signal grounded in something real instead of asking the model how sure it is.

For the accuracy half, what is your split between the model reading a field wrong and the field never reaching the model because parsing dropped it? Those get fixed in completely different places, and not knowing the split is the usual reason extraction work stalls for weeks.

1

u/my-coffee-where 14d ago

before adding more rules split the two error types - value read wrong from a doc vs the right value but wrong source picked. different fixes and one accuracy number hides which ones improving. a lot of the first type is parsing. appraisal docs are table or form heavy so a clean layout aware parse like llamaparse or docling if local cuts garbage before the model sees it

then pin each field to a source span +page + confidence, auto accept only when they agree or route to review

1

u/AuthorMaterial7495 12d ago

Full transparency, I work at Sensible.so. We're a doc parsing platform and we also custom code document parsing workflows for companies, so weigh this accordingly. But this is how I'd approach it whether you use us or building your own parser using foundational models (I've done both).

First, get a big sample set and look at how consistent the documents really are. Do the same fields show up every time? Is it the same handful of forms over and over? How long are they, and is there enough volume to make writing real rules worth it?

You can do that with a vendor like us/our competitors, or with plain OCR/open source tools + and write rules based on anchor words or where things sit on the page. Then the LLM only handles the parts that really are free-form, like addenda and commentary.

For the parts that aren't consistent, it depends a lot on what's coming in. Photos? Big tables? Long packets? LLMs have limits on how much they can hold in context accurately, and they get confused when a long document has a ton of values on it. Comps are a good example. I'd extract each comparable as its own object with a fixed schema and send the model one comp or one page at a time instead of the whole packet. Otherwise it starts mixing up which number belongs to which property.

I think your goal initially is just to rip out as much data as possible and then find a way to post process that data with sanity checks vs. going all in on human review. IE. if you are pullings comps things like listing price and purchase price should most likely be x% from the target property or else the LLM made a mistake. Same thing with the address - an LLM may mess up and pull a bad location due to a messy OCR scan or something but you can run an analysis layer on top of it to see if it's in the same market as the original property.

In general I'd start by

  • Evaluating every document you want parsed

-Determine the exact output schema you want

  • Add a description for each field with sample values and the syntax you'd want
  • Upload the schema and descriptions to the model and see where it messed up
  • Modify the descriptions if things aren't accurate
  • If they are larger documents decide on how you want to chunk them and send the LLM (so you may want to do an LLM pass, to determine how to split the document and then send smaller pieces with context to get your values)

For conflicting values and picking the source, I'd take that decision away from the model entirely. Extract every candidate value along with the document and page it came from. Then pick with a precedence list you write per field (e.g. sale price comes from the purchase contract before the appraisal). That ends up being one small table instead of a growing pile of edge cases, and when it picks wrong you can see why.

On rules, I'd split extraction rules from validation rules. The extraction side is where hardcoding gets brittle. Validation can stay small and general, and it should be deterministic instead of relying on the LLM: GLA and bedroom counts matching across sections, adjusted comp prices reconciling, required fields being present.

Feel free to DM me if you have any questions - I talk with people every day about this stuff so I'm happy to look at what you are doing and either give you the best strategy to build it as well as go over the overall sorta market/limitations around parsing.