datalab-to/omni_extract_bench
Omni Extract Bench We weren’t satisfied with the current benchmarking options for extraction. They were biased, didn’t use realistic data and were hard to audit. Our view is that an extraction benchmark should do two things: Help customers choose the right vendor; and Give engineers a way to diagnose what’s actually going wrong in a given model. That’s why we built OmniExtractBench. OmniExtractBench is a comprehensive structured extraction benchmark, developed by Datalab.… See the full description on the dataset page: https://huggingface.co/datasets/datalab-to/omni_extract_bench.
Add 40-document subset manifest (predictive across vendors, 2026-09-22)
Update llamaextract figures
Add blog post link
Update README: install/benchmark/score/predict sections, add usage gif and docs
Link the GitHub repo from the intro
Rewrite the card, and add the figures it references
Licence the collection CC BY 4.0, with per-subset terms
Replace with the 620-document scored set (part 5)
Replace with the 620-document scored set (part 4)
Replace with the 620-document scored set (part 3)
Replace with the 620-document scored set (part 2)
Replace with the 620-document scored set
Revert rename: Omni Extract Bench
Rename to Fair Extract Bench
Headline is the mean over all documents
Ship all 660 scored documents (was 169); defective pairs moved out of data/; card states UNIFIED headline and run protocol
Add internal licence text
Add longarray licence text
Remove gt_corrections.json: it was the REJECTED candidate pile, never applied
Record the single applied correction as provenance, not an overlay
Ground truth is as-shipped; remove the mislabelled corrections file
Frozen 169-document evaluation set with verified GT corrections
initial commit
