llamaindex/ExtractBench
ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.
3220k
1version https://git-lfs.github.com/spec/v12oid sha256:d60c73504d84a6439e13b1465103ed43a28e881f6fb6b845812155501464d0443size 122073784 