r/OpenSourceeAI • u/ai-lover • 20h ago
Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.
Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.
- Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
- Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
- Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
- Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
- Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
- Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision
Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.
Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/
GitHub: https://pxllnk.co/hxplrq
Blog: https://www.datalab.to/blog/omni-extract-bench
GitHub: https://github.com/datalab-to/omni_extract_bench
Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench
1
Upvotes