datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
documentation-images
This dataset contains images used in the documentation of HuggingFace's libraries.
HF Team: Please make sure you optimize the assets before uploading them.
My favorite tool for this is https://tinypng.com/.
documentation-imagesdocument-haystack
Document Haystack Dataset
This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”.
📑 Abstract Paper
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.wikitext_document_level
Wikitext Document Level
This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below.
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.documentation-imagesdocumentation-mediadocumentation-imagesdocumentation-imagesdocumentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library.
foia-reading-room-documents
Foia Reading Room Documents
Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is of the
original bytes.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.JFK-Assassination-Records-2025-Documents-Releasedocumentation-imagesdocumentation-imagesDocumentVQAdocument-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.apex-r1-real-world-documents
Apex-R1 Real-World Benchmark Documents
This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation.
The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks.
Contents
benchmark_documents/
EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.documentation-imageslocus-pretrain-documents-public
Locus pretraining: documents public
Setup only. No corpus release has been published and no canary has been run for this format.
This repository will hold versioned indexed token artifacts. Tokens are little-endian int32 binary arrays; signed int64 offsets identify complete documents or unpadded pieces. Loss masks are bit-packed, with Parquet indexes and provenance sidecars. Training packs pieces using RSDB.
Only a complete, verified manifest pinned to an immutable commit… See the full description on the dataset page: https://huggingface.co/datasets/sunnysetia/locus-pretrain-documents-public.AI2_Alphabot_2_stamp_document
AI2_Alphabot_2_stamp_document
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 987
Total Frames: 369702
FPS: 30
Dataset Size: 7.18 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb,
cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stamp_document.jurisdb-legal-documents
JurisDB - Brazilian Legal Documents Dataset
Dataset Description
This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU).
Dataset Structure
.
├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/
│ ├── leis_estaduais/
│ ├── leis_federais/
│ └── ...
└── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/
├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.Embrapa-ai-documents-markdownform_understanding_in_noisy_scanned_documents_plus
Dataset Card for Form Understanding in Noisy Scanned Documents Plus
This is a FiftyOne dataset with 1026 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.document_alignment_dataset-Sinhala-Tamil-English
Dataset summary
This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages.
Data had been crawled from the following news websites.
News Source
url
Army
https://www.army.lk/
Hiru
http://www.hirunews.lk
ITN
https://www.newsfirst.lk
Newsfirst
https://www.itnnews.lk
The aligned documents have been manually annotated.
Dataset
The folder structure for each news source is as follows.
army
|--Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English.document-haystack-10pages
Dataset Card for document-haystack-10pages
This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/document-haystack-10pages")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.sam-solicitation-documents
Sam Solicitation Documents
Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.documenti-societari-italiani-rag-evaldocumentation-imagesArabic_Documents_Dataset_PDF
Arabic Documents Dataset (PDF)
This dataset contains a collection of Arabic-language documents in PDF format. The corpus includes books, articles, reports, and educational materials written in Modern Standard Arabic and regional variants. It is curated to support AI research in document understanding, Arabic OCR, and text extraction from complex layouts.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Arabic_Documents_Dataset_PDF.Agilex_Cobot_Magic_zip_up_the_document_bag
Agilex_Cobot_Magic_zip_up_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pull
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.French_Documents_Dataset_PDF
French Documents Dataset (PDF)
This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.
