datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.irs-formssynthetic-us-forms-preview
SymageDocs — Synthetic US Forms Preview
A small, CC-BY-4.0, fully synthetic document-AI training set: 525
labeled page images across six families of US business and government forms,
each page shipping FUNSD ground truth plus a LayoutLM-ready token/bbox/tag view.
This is a preview subset. It exists so you can load real output from the
SymageDocs generator, inspect the label quality, and
decide whether generating your own corpus is worth your time — without an
account, an email… See the full description on the dataset page: https://huggingface.co/datasets/Symage/synthetic-us-forms-preview.handwriting_forms
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ift/handwriting_forms.US_tax_forms_donut
NIST-SD2 (US Tax Forms) Donut Dataset
Processed version of NIST Special Database 2 for document understanding tasks, formatted for use with the Donut architecture. Contains 5,590 annotated document images (5,031 train, 559 test) across 20 tax form classes.
Description
Curated by: National Institute of Standards and Technology (NIST)
License: MIT
Total Size: 947.33 MB
Annotations:
Class labels (20 tax form types)
Full text ground truth
legal-templates-multilingual
Forms Legal — Multilingual Legal Templates Corpus
19,419 legal document templates across 36 jurisdictions in 12 languages, expanded to 22,876 (document × locale) rows. Released under CC-BY-4.0 by forms-legal.com.
Quick description
A multilingual corpus of structured legal document templates spanning 36 jurisdictions. Each document includes a multi-section editorial brief (whatIs, whenNeeded, keyElements, howToFill, legalRequirements, commonMistakes), 5-8… See the full description on the dataset page: https://huggingface.co/datasets/forms-legal/legal-templates-multilingual.handwriting_forms_cleaned
handwriting_forms_cleaned
The handwriting_forms__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
1,360
QA turns
5,198
answers rewritten by the cleaning pass
166
QA created by the cleaning pass (new_qa)
3,860 (74.3%)
shards
1
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/handwriting_forms_cleaned.OMR-forms
Dataset Card for "OMR-forms"
More Information needed
coherent-forms-1040-cms1500-i9
SymageDocs — Coherent US Tax / Health / Employment Forms (FUNSD)
A fully synthetic document-AI training set: three US forms — IRS Form 1040,
CMS-1500, and USCIS Form I-9 — filled from the same synthetic
identity, so name / SSN / address / employer flow consistently across all
three renderings. Each page ships with FUNSD ground truth (word boxes,
entity labels, key–value linking) plus a LayoutLMv3-ready token/bbox/tag view.
3,000 page-level image + annotation rows (train 2,400 /… See the full description on the dataset page: https://huggingface.co/datasets/Symage/coherent-forms-1040-cms1500-i9.uk-benefit-forms-structured
UK Benefit Forms Structured Dataset
A structured dataset of 120 UK government benefit and legal forms, extracted and processed for use in AI-assisted form-filling applications. Built as part of the EasyClaimAI project.
Why This Dataset Exists
Millions of people in the UK struggle with complex government forms — benefit claims, legal applications, pension forms. The language is dense, the guidance is buried, and mistakes can cost people money or delay vital support.
This… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/uk-benefit-forms-structured.handwriting_forms_cleaned
handwriting_forms_cleaned
The handwriting_forms__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
1,360
QA turns
5,198
answers rewritten by the cleaning pass
166
QA created by the cleaning pass (new_qa)
3,860 (74.3%)
shards
1
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/handwriting_forms_cleaned.medical-forms-datasetqa-formsSanskrit-verb-forms
Sanskrit Verb Forms Dataset (संस्कृत धातु रूप संग्रह)
Overview
description: |
A comprehensive dataset containing Sanskrit verb conjugations (dhatu roop) with 10,348 unique entries.
Each entry provides the complete information about a Sanskrit verb form, including:
धातु (Dhatu): The root verb
पद (Pada): Voice of the verb (परस्मैपद/आत्मनेपद)
लकार (Lakara): Tense/mood of the verb
पुरुष (Purusha): Person (प्रथम/मध्यम/उत्तम)
वचन (Vachana): Number (एकवचन/द्विवचन/बहुवचन)… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-verb-forms.arabic-legal-documents-and-verified-formsafrica-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms
Youth working-age population by sex, age and forms of transition (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms.asia-owid-forms-of-homelessness-included-in-available-statistics
Forms Of Homelessness Included In Available Statistics | Asia (Our World in Data)
🌏 33 observations · 33 Asia countries · 2010–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 33 observations of Forms Of Homelessness Included In Available Statistics data across 33 Asia countries, spanning 2010–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Forms Of Homelessness… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-forms-of-homelessness-included-in-available-statistics.africa-uganda-experience-of-various-forms-of-violence-801c079c
Experience of Various Forms of Violence | Africa (Uganda Bureau of Statistics)
7 rows - 1 Africa country/area - 2025 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 7 rows from Uganda Bureau of Statistics, covering Experience of Various Forms of Violence. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-experience-of-various-forms-of-violence-801c079c.europe-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms
Youth working-age population by sex, age and forms of transition (thousands) | Europe (ILOSTAT)
🇪🇺 89,895 observations · 38 Europe countries · 1991–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 89,895 observations of Population data across 38 Europe countries, spanning 1991–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms.africa-uganda-forms-of-spousal-violence-4ca2303e
Forms of Spousal Violence | Africa (Uganda Bureau of Statistics)
8 rows - 1 Africa country/area - 2025 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 8 rows from Uganda Bureau of Statistics, covering Forms of Spousal Violence. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Official statistics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-forms-of-spousal-violence-4ca2303e.tune-forms
Dataset Card for "tune-forms"
More Information needed
donut-medical-formsafrica-owid-forms-of-homelessness-included-in-available-statistics
Forms Of Homelessness Included In Available Statistics | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-forms-of-homelessness-included-in-available-statistics.fixed_forms
Dataset Card for "fixed_forms"
We propose a new dataset pf fixed poetic forms in English, which can be used both for literary analysis and for training of poetry generators (including large language models).
The general structure of the dataset is as follows: it contains 12 rows - according to the number of fixed forms, which are:
ballade
rondeau
triolet
ottava rima
italian (petrarchan) sonnet
french sonnet
english (shakespearean) sonnet
ode stanza
elegiac distich (couplet)
haiku… See the full description on the dataset page: https://huggingface.co/datasets/avfattakhova/fixed_forms.formsasia-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms
Youth working-age population by sex, age and forms of transition (thousands) | Asia (ILOSTAT)
🌏 10,399 observations · 17 Asia countries · 1999–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 10,399 observations of Population data across 17 Asia countries, spanning 1999–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms.europe-owid-forms-of-homelessness-included-in-available-statistics
Forms Of Homelessness Included In Available Statistics | Europe (Our World in Data)
🇪🇺 39 observations · 39 Europe countries · 2011–2024 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 39 observations of Forms Of Homelessness Included In Available Statistics data across 39 Europe countries, spanning 2011–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Forms Of… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-forms-of-homelessness-included-in-available-statistics.
