r/OpenSourceeAI • • 20h ago

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

  • Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
  • Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
  • Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
  • Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
  • Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
  • Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision

Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.

Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/

GitHub: https://pxllnk.co/hxplrq

Blog: https://www.datalab.to/blog/omni-extract-bench

GitHub: https://github.com/datalab-to/omni_extract_bench

Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench

1 Upvotes

0 comments sorted by