CoolFace
Datasetpublic

ibm-research/VAREX

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.

sourceHugging Facecdla-permissive-2.0updated 6mo agoView on Hugging Face
7likes1.9kdownloads
8 commits on main
2dfc3386mo ago

Update README.md

Udibarzi
ff36b6c6mo ago

Update README.md

Udibarzi
b1c05406mo ago

Upload 1777 files

Udibarzi
df976966mo ago

Upload 4 files

Udibarzi
f7fca2b6mo ago

Upload manifest.json with huggingface_hub

Udibarzi
ec5f7c86mo ago

Upload field_exclusions.json with huggingface_hub

Udibarzi
ca17a476mo ago

Upload README.md with huggingface_hub

Udibarzi
162f9356mo ago

initial commit

Udibarzi