llamaindex/ExtractBench
ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.
3220k
1version https://git-lfs.github.com/spec/v12oid sha256:7579253d09326763598ea3c2592f980f1119784d7bcd21979bbfe43b3bcbee393size 2348664 