datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.Agilex_Cobot_Magic_zip_up_the_document_bag
Agilex_Cobot_Magic_zip_up_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pull
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.cc100-documents
cc100-documents
This dataset is a restructured version of the CC-100 (statmt/cc100) dataset.
In the original dataset, each instance corresponds to a single paragraph (or a document boundary).
In this version, the data has been reformed so that each instance corresponds to a single, complete document.
This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library.
Languages
The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.RMC-AIDA-L_organise_the_document_bag
RMC-AIDA-L_organise_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
pull
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.Chinese_Debate_Documents
Dataset Card for Chinese Debate Documents
ASR-transcribed corpus of competitive Mandarin Chinese university debates,
speaker-segmented and timestamped, with topic / round / team metadata parsed
from the source filenames.
Loading the Dataset
from datasets import load_dataset
ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train")
print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"])
for seg in ds[0]["segments"][:3]:
print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.data_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.named_entity_recognition_document_contextbag-of-documents
Bag-of-Documents: Product Search Dataset
Blog post: Distilling Retrieval Pipelines to a Single Embedding Model
Live demo: huggingface.co/spaces/dtunkelang/bag-of-documents-demo
Code: github.com/dtunkelang/bag-of-documents
Dataset Description
A large-scale bag-of-documents dataset for e-commerce product search, built on Amazon product data. Each search query is represented as a distribution of relevant products in embedding space, captured by a centroid vector and a… See the full description on the dataset page: https://huggingface.co/datasets/dtunkelang/bag-of-documents.dta-documents
Deutsches Textarchiv (DTA) Documents
This datasets hosts all documents from the Deutsches Textarchiv (DTA).
One row per work of the Deutsches Textarchiv (DTA), built from the
official DTA dump dta_komplett_2026-02-10 (TCF format). The dataset contains 5,480 documents spanning
1472 to 1987 with about 205M tokens and 1.4B characters of German text, with two text views per document:
text: the historical, layout-faithful transcription (line breaks, long s ſ, combining diacritics… See the full description on the dataset page: https://huggingface.co/datasets/histde/dta-documents.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.Turkish-Law-Documents-700k-clustered
Turkish Legal Documents Clustering Dataset
A comprehensive dataset of 700,000 Turkish legal documents from the two primary sources of legal precedent in Turkey, clustered using multiple emebdding models and algorithms to enable research, analysis, and machine learning applications.
Overview
This repository contains a large-scale document clustering pipeline and dataset for Turkish legal documents sourced from:
Yargıtay - Turkey's highest court of appeal for civil and… See the full description on the dataset page: https://huggingface.co/datasets/erdem-erdem/Turkish-Law-Documents-700k-clustered.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.dutch-legal-documents
Dutch Legal Documents
A comprehensive collection of 1.14 million Dutch legal documents, including court rulings, parliamentary documents, and EU legislation.
Sources
Source
Type
Documents
Description
Rechtspraak.nl
Court rulings
882,212
All Dutch court rulings from Hoge Raad, Raad van State, Gerechtshoven, Rechtbanken, CRvB, CBb
Officiële Bekendmakingen
Parliamentary docs
255,837
Kamerstukken, Handelingen, Memories van Toelichting, Staatsblad
EUR-Lex
EU… See the full description on the dataset page: https://huggingface.co/datasets/sheggle/dutch-legal-documents.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm..
Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.all-document-text-data
Climate Policy Radar Open Data
This repo contains the full text data of all of the documents from the Climate Policy Radar database (CPR), which is also available at Climate Change Laws of the World (CCLW).
Please note that this replaces the Global Stocktake open dataset: that data, including all NDCs and IPCC reports is now a subset of this dataset.
What’s in this dataset
This dataset contains two corpus types (groups of the same types or sources of documents) which… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/all-document-text-data.doclaynet-document-level
DocLayNet Document-Level Reconstruction and 8K Expansion
This dataset is a normalized, one-row-per-document view over the page-level
DocLayNet v1.1
dataset. Pages are grouped using DocLayNet's source metadata and ordered by their original
page number.
Dataset summary
2,944 logical documents
80,863 observed pages
896 complete document groups
2,048 partial document groups
Train: 2,355 documents / 60,810 pages
Validation: 294 documents / 7,964 pages
Test: 295… See the full description on the dataset page: https://huggingface.co/datasets/operant-ai/doclaynet-document-level.document-corpus-v3-open
Document Corpus v3 Open
document-corpus-v3-open is the redistribution-compatible slice of the exact
byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT
experiments. It contains 869,739 filtered documents and
2.192 GB of UTF-8 text before Parquet compression.
This is not the complete internal document-corpus-v3. Restricted,
unknown-license, and share-alike sources were excluded conservatively. Every
included row comes from an upstream dataset whose card… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.dodis-historical-documents
Dataset Overview
The Swiss Historical Archive of DOdis Works (SHADOW) is a large-scale benchmark dataset derived from the resources of Dodis, an independent research center dedicated to the study of Swiss foreign policy and Switzerland’s international relations. Dodis has curated and published nearly 50,000 key documents that document the administrative practices and decision-making processes of the Swiss federal administration.
The underlying corpus consists of notes, letters… See the full description on the dataset page: https://huggingface.co/datasets/prg-unibe/dodis-historical-documents.global-stocktake-documents
Global Stocktake Open Data
This repo contains the data for the first UNFCCC Global Stocktake. The data consists of document metadata from sources relevant to the Global Stocktake process, as well as full text parsed from the majority of the documents.
The files in this dataset are as follows:
metadata.csv: a CSV containing document metadata for each document we have collected. This metadata may not be the same as what's stored in the source databases – we have cleaned and added… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/global-stocktake-documents.SO_KGXQR_DOCUMENT
Dataset Card for "SO_KGXQR_DOCUMENT"
More Information needed
aws_bedrock_documentation_demo
Aws Bedrock Documentation Demo
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/aws_bedrock_documentation_demo.open-scholarly-document-catalog-10k
Search interactively
·
Product and methodology
·
Full catalog
Rights-Audited Scholarly PDF Catalog: 10K Validated Sample
Evaluation sample / inspect before you buy. Metadata and source links
only. No PDFs or extracted full text are redistributed.
Build scientific-literature search, RAG, discovery, and corpus-acquisition
workflows from scholarly metadata with current PDF-link evidence and per-record
Creative Commons or NASA rights… See the full description on the dataset page: https://huggingface.co/datasets/PlethoraSolutions/open-scholarly-document-catalog-10k.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/abby2231/indian-legal-documents.research-document-archive
Research Document Archive
Computational analysis of 234,630 declassified U.S. government documents across 7 archival collections. Output of a 13-step ML pipeline extracting OCR text, entities, topics, keywords, redactions, and semantic embeddings from 3.1 million pages.
Files
File
Rows
Description
documents.parquet
234,630
Document metadata: id, source_section, file_path, file_hash, total_pages
pages/<section>.parquet
3.1M
Per-page OCR text + 1536-dim… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/research-document-archive.muninn-ww1-documents
Muninn WWI Documents (CEF Attestation Papers & War Diaries)
A tabular conversion of the document records in the Muninn Project's
World War I Linked Open Data archive (https://rdf.muninn-project.org/). The Muninn Project is a
research project that extracts structured data from digitized WWI-era archival documents.
The dataset covers 28,742 documents — 28,718 Canadian Expeditionary Force (CEF) attestation
papers
(enlistment forms of individual soldiers) and 24 Canadian unit war… See the full description on the dataset page: https://huggingface.co/datasets/biglam/muninn-ww1-documents.document-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.bioreason-pro-sft-reasoning-documents
BioReason-Pro SFT Reasoning Documents
Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training.
The upstream dataset wanglab/bioreason-pro-sft-reasoning-data
ships the assistant side of each training example (reasoning, final_answer) alongside the raw
biological context columns, but not the assembled prompt. The prompt cannot be recovered from the
data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.california-housing-documented
