CoolFace
Datasetpublic

llamaindex/ExtractBench

ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
32likes20kdownloads
long.jsonl4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:ac5b22bbf810548a5e086e792d4afe0aa98dc572d24deceab820a2558570f3bd3size 1599657894