
Datalab OmniExtractBench Scores PDF-to-JSON Extraction by Value
Published by AINave Editorial
Datalab’s OmniExtractBench tests a practical failure mode in document automation: whether a system can fill a JSON schema from a PDF without dropping or making up values. It combines 620 documents from four existing benchmark sets and applies one deterministic scorer, which assigns verdicts to individual values rather than returning only a document-level score. The launch coverage was sponsored by Datalab, and the results it reports are Datalab’s evaluation, not independent validation. The benchmark and its reported results
One corpus, several kinds of documents
The collection combines 329 ExtractBench documents from LlamaIndex, 202 synthetic documents from Datalab, 47 LongExtractBench documents from micro1, and 42 LongArray-Extract documents from Extend. The source describes a mix of forms, filings, decks, dense scalar schemas and large tables. That variety matters because a benchmark dominated by one document shape can conceal weaknesses on another. The corpus breakdown and content
The corpus also ranges from short files to long ones: 128 documents are a single page, while 33 documents exceed 100 pages and account for 40% of the pages in the collection. This gives teams more than a test of clean, short forms, although it does not establish how any particular system will perform on a team’s own documents. Document-length figures
The scorer explains how extraction failed
OmniExtractBench flattens predicted and reference JSON into individual value paths, normalizes values, and compares them. For tables, it pairs rows by their contents using the Hungarian algorithm rather than relying only on row position. In the launch article’s example, a 100-row table missing its first row scored 0% under positional comparison and 99% with OmniExtractBench’s scorer. That example illustrates why alignment matters: one omitted row can otherwise make every following row appear wrong. Scoring method and example
Each value receives one of six verdicts: matched, misread, unfound, fabricated, invented item or invented field. The scorer uses these verdicts to calculate accuracy, precision and recall, which reveal different error patterns. Empty strings, None and whitespace count as omissions; values such as “NA” and “-” remain answers. This rule is intended to stop systems from gaining credit by filling optional fields with empty values. Verdict definitions and null handling
Rankings are less useful without error profiles
In Datalab’s reported evaluation of 10 system configurations, Datalab’s accurate mode scored 93.85 accuracy. Datalab balanced scored 93.48, and Reducto deepextract v2 scored 93.47. The close scores make the error breakdown more informative than a simple ranking: the article reports GPT 5.6-sol had higher precision than recall, while LlamaExtract had higher recall than precision. In plain terms, those patterns distinguish a tendency to leave fields out from one to return extra or unsupported values. [Reported scores and error profiles](https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/?utmsource=flipboard&utm_content=Marktechpost%2Fmagazine%2FAI+Research+News)
The scorer is described as available through PyPI under Apache 2.0, with the dataset under CC BY 4.0. Running vendor systems again requires users’ own API keys and paid credits, so an open scoring tool does not make the entire evaluation cost-free. And because the launch coverage was sponsored by Datalab, the reported leaderboard is best read as a useful, inspectable comparison from the benchmark’s creator, not a neutral independent verdict. Licensing, rerun requirements and sponsorship disclosure
The benchmark’s clearest contribution is not a definitive winner. It is a way to inspect what a score means: whether the system matched the source, missed a field or supplied something the document did not support. That distinction is especially useful when two systems land near the same overall accuracy but fail in different ways. Scorer design and reported comparisons





















