ibm-research/VAREX
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.
Update README.md
Update README.md
Upload 1777 files
Upload 4 files
Upload manifest.json with huggingface_hub
Upload field_exclusions.json with huggingface_hub
Upload README.md with huggingface_hub
initial commit
