forms
Datasets
All datasets matching “forms”FormStruct-Bench
FormStruct-Bench
Dataset Description
FormStruct-Bench is a multilingual benchmark for extracting the semantic and
spatial structure of forms from document images. The repository combines a
7,000-page main benchmark, a controlled visual-degradation set, and
template-level layout annotations. It supports evaluation of vision-language
models and document AI systems on hierarchical key-value extraction, document
structure recovery, region localization, table and… See the full description on the dataset page: https://huggingface.co/datasets/D2I-CUHK-Shenzhen/FormStruct-Bench.cua-s1-forms
cua-s1-forms (dataset)
Synthetic + real training/eval data for cua-ai/cua-s1-forms,
a jev-like one-pass option scorer for GUI form filling behind
cua-driver.
Generator source: cua_s1/synth.py in
https://github.com/trycua/cua/tree/main/libs/cua-s1.
Files
file
rows
source
train.jsonl
~150k
synthetic
validation.jsonl
~18k
synthetic
test.jsonl
~20k
synthetic, form-signature-disjoint from train/validation
demo.jsonl
196
real: 3 real JevBrowser form… See the full description on the dataset page: https://huggingface.co/datasets/cua-ai/cua-s1-forms.Intel-WebCorpus-formsirs-formsANC-formsindian_dance_formsThis dataset is taken from https://www.kaggle.com/datasets/aditya48/indian-dance-form-classification but is originally from the Hackerearth deep learning contest of identifying Indian dance forms. All the credits of dataset goes to them.
Content
The dataset consists of 599 images belonging to 8 categories, namely manipuri, bharatanatyam, odissi, kathakali, kathak, sattriya, kuchipudi, and mohiniyattam. The original dataset was quite unstructured and all the images were put together.… See the full description on the dataset page: https://huggingface.co/datasets/tanmaykm/indian_dance_forms.
