datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.docmathdoc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
DocDownstream-2.0DocDownstream 2.0 is a collection of MP-DocVQA, DUDE, NewsVideoQA used in DocOwl2.
For MP-DocVQA and DUDE, the maximum number of pages of each sample is set to 20.
For NewsVideoQA, the maximum number of frames of each sample is set to 20.
XL-DocBench
XL-DocBench
Evidence-grounded reasoning across hundreds or thousands of pages.
Fully verified by 194 human experts.
Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡,
Bei Liu2,*, Yifan Yang2, Qi Dai2,
Ruichun Ma2, Kai Qiu2, Yunsheng Li2,
Dongdong Chen2, Chong Luo2,
Zhenzhong Chen1, Baining Guo2
1Wuhan University 2Microsoft
†Equal contribution ‡Work done during an internship at MSRA
*Project leader
Project Page ·
Paper ·
Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.DocFinQAmarkdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.znanio-documents
Dataset Card for Znanio.ru Educational Documents
Dataset Summary
This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.Re-DocRED
Re-DocRED Dataset
This repository contains the dataset of our EMNLP 2022 research paper Revisiting DocRED – Addressing the False Negative Problem
in Relation Extraction.
DocRED is a widely used benchmark for document-level relation extraction. However, the DocRED dataset contains a significant percentage of false negative examples (incomplete annotation). We revised 4,053 documents in the DocRED dataset and resolved its problems. We released this dataset as: Re-DocRED dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/tonytan48/Re-DocRED.long-doc_book_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: dev
us-caselaw
US Caselaw
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
The American case-law record in one dataset — published opinions from the Free Law Project /
CourtListener bulk export of 2026-06-30: full text, all opinion types (majority, concurring,
dissenting, per curiam… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw.dockersv1text-code-galeras-code-generation-from-docstring-3k-dedupedDocTalk
📁 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
➤ 📖 Paper Link
DocTalk is a large-scale, synthetic dialogue corpus created via a three-stage pipeline that converts clusters of related Wikipedia documents into multi-turn, multi-topic information-seeking conversations.
The pipeline comprises:
Document Graph Construction: Sampling up to three related Wikipedia articles per anchor document via a weighted random walk on a directed… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/DocTalk.docker-build-cache-trajectories
Docker Build Cache Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/docker-build-cache-trajectories.docsmil-docs
What is this?
A curated selection of manuals and documents from the US military and other departments. All data was manually scraped from publicly available sources.
The PDF's and EPUB files were converted to markdown using the amazing Marker github repository by Vik Paruchuri.
Sources:
United States Army Central Army Repository
Marines Publications
Federation of American Scientists Intelligence Resource Program
en-document-classification
English Document Classification Dataset
This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora.
Dataset Summary
The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.document-review-source1k
New 1K title extraction corpus
Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified.
1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.LoCoV1-Documents
LoCoV1 Documents
The documents for the LoCoV1 dataset from the paper, "Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT"
How to Use
To load the dataset, use the following command:
from datasets import load_dataset
dataset = load_dataset("hazyresearch/LoCoV1-Documents")
To load a specific subset, such as SummScreenFD, use the following command:
from datasets import load_dataset
dataset = load_dataset("hazyresearch/LoCoV1-Documents")
def… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/LoCoV1-Documents.texas-caselaw
Texas Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of
2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/texas-caselaw.document-qna-chroma-anyscale-logsus-regulations
US Federal Regulations — the CFR, held word for word
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
222,767 regulations — 219,061 CFR sections and 3,706 appendices — across all 49 titles
of the Code of Federal Regulations, each one as the agency publishes it.
Statutes say… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-regulations.us-caselaw-ca
California Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Full text of 242,989 California appellate opinion documents from the public record. The base
slice is 242,831 documents from the Free Law Project / CourtListener bulk export of 2026-06-30.… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-ca.godot_4_docsDataset generated for Godot 4 docs using Glaive.
us-statutes
US Statutes — held word for word
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
1,295,620 statute sections: the whole United States Code (60,433 sections, all 53 titles)
plus 27 states (1,235,187 sections), each section as its legislature publishes it, with
the URL it was… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-statutes.tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.sorrel-T-qwen3-8b-base-seed0-documentsus-caselaw-wi
Wisconsin Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Full text of 72,350 Wisconsin appellate opinion documents from the public record,
sliced from the Free Law Project / CourtListener bulk export of 2026-06-30.
Court coverage (2 court ids, explicit allowlist —… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-wi.
