datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FormStruct-Bench
FormStruct-Bench
Dataset Description
FormStruct-Bench is a multilingual benchmark for extracting the semantic and
spatial structure of forms from document images. The repository combines a
7,000-page main benchmark, a controlled visual-degradation set, and
template-level layout annotations. It supports evaluation of vision-language
models and document AI systems on hierarchical key-value extraction, document
structure recovery, region localization, table and… See the full description on the dataset page: https://huggingface.co/datasets/D2I-CUHK-Shenzhen/FormStruct-Bench.cua-s1-forms
cua-s1-forms (dataset)
Synthetic + real training/eval data for cua-ai/cua-s1-forms,
a jev-like one-pass option scorer for GUI form filling behind
cua-driver.
Generator source: cua_s1/synth.py in
https://github.com/trycua/cua/tree/main/libs/cua-s1.
Files
file
rows
source
train.jsonl
~150k
synthetic
validation.jsonl
~18k
synthetic
test.jsonl
~20k
synthetic, form-signature-disjoint from train/validation
demo.jsonl
196
real: 3 real JevBrowser form… See the full description on the dataset page: https://huggingface.co/datasets/cua-ai/cua-s1-forms.hebrew_suffix_verbal_forms
Suffixed Verbal Forms Detection Dataset for Modern Hebrew
Dataset Summary
This dataset contains annotated Hebrew sentences containing verbal forms that are ambiguous as to whether they include a pronominal suffix or not (e.g., the Hebrew word lamed-yod-mem-daled-vav can be understood as either "he taught him" or "they taught"). The goal of the dataset is to support tasks involving the identification and disambiguation of verbs with pronominal suffixes in Hebrew literature… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/hebrew_suffix_verbal_forms.theyyam-forms-kbPDF-FORMSforms-dataelection-formss1_forms_benchforms-filling-r1-distilldatetime-alpaca-formstone-pass-sv-forms-synthetic
One-Pass SV-Forms synthetic corpus (Swedish)
This dataset is entirely synthetic. It was generated by a script from a concept catalogue, not
collected from people or customer submissions. Names and organisations are constructed,
email addresses use .invalid, and identifier-shaped values are generated locally.
They are not checked against registries; coincidental matches with real names or
identifiers cannot be ruled out. The dataset is published so the recipe behind… See the full description on the dataset page: https://huggingface.co/datasets/precisit/one-pass-sv-forms-synthetic.
