CoolFace
Datasetpublic

ibm-research/VAREX

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.

sourceHugging Facecdla-permissive-2.0updated 6mo agoView on Hugging Face
7likes1.9kdownloads

ibm-research/VAREX · main · files are served by the source, never re-hosted here