datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.markdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.znanio-documents
Dataset Card for Znanio.ru Educational Documents
Dataset Summary
This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.en-document-classification
English Document Classification Dataset
This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora.
Dataset Summary
The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.LoCoV1-Documents
LoCoV1 Documents
The documents for the LoCoV1 dataset from the paper, "Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT"
How to Use
To load the dataset, use the following command:
from datasets import load_dataset
dataset = load_dataset("hazyresearch/LoCoV1-Documents")
To load a specific subset, such as SummScreenFD, use the following command:
from datasets import load_dataset
dataset = load_dataset("hazyresearch/LoCoV1-Documents")
def… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/LoCoV1-Documents.document-qna-chroma-anyscale-logssorrel-T-qwen3-8b-base-seed0-documentssorrel-T-qwen3-1.7b-base-seed0-documentsbag-of-documents
Bag-of-Documents: Product Search Dataset
Blog post: Distilling Retrieval Pipelines to a Single Embedding Model
Live demo: huggingface.co/spaces/dtunkelang/bag-of-documents-demo
Code: github.com/dtunkelang/bag-of-documents
Dataset Description
A large-scale bag-of-documents dataset for e-commerce product search, built on Amazon product data. Each search query is represented as a distribution of relevant products in embedding space, captured by a centroid vector and a… See the full description on the dataset page: https://huggingface.co/datasets/dtunkelang/bag-of-documents.multi-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.certified-document-qa
Certified Document QA: span-verified, absence-aware
79,400+ rows · every claim machine-re-checkable · zero
frontier-model-derived tokens · includes filings newer than every major
training cutoff · plus a free 127K-token verified long-context task-set.
Of the 79,438 published rows, 7,921 carry an inline machine-checkable
certificate column (needle_public, needle_expansion_v120,
absence_public, multihop_public, both teasers, and the dated multihop
splits). A further 54,837 rows —… See the full description on the dataset page: https://huggingface.co/datasets/SovNodeAI/certified-document-qa.legal_case_document_summarization
Dataset Card for LegalCaseDocumentSummarization
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/legal_case_document_summarization.en-document-format-classification
English Document Format Classification Dataset
English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats.
Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks.
Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage).
Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.en-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.document-level-word-alignment
Document-Level Word Alignment
Document-level word alignment data for six language pairs — English–French (en-fr), English–Romanian (en-ro), English–Japanese (en-ja), English–Chinese (en-zh), Latin–Ancient Greek (la-gr), and English–Czech (en-cz) — reconstructed from existing sentence-level, human-annotated word alignment gold standards.
Document-level examples are built by grouping sentence-level annotations by document membership and sentence order; the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/document-level-word-alignment.sorrel-T-qwen3-14b-base-seed0-documentsworldbank-project-documents
Dataset Card for World Bank Project Documents
Dataset Summary
This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes
the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed
by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets.
Supported Tasks and Leaderboards
No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.documentssorrel-T-mistral-small-24b-base-seed0-documentssorrel-T-olmo-2-32b-seed0-documentslegal-ai-regulation-documents
SafeLegalAI Legal AI Regulation Documents (versioned)
Which official texts govern AI in legal practice, what do they require, and how have they changed?
515 documents · 560 versions tracked · 17 jurisdictions · last checked 2026-09-06 · synced from safelegalai.com on 2026-09-08.
Every court rule, practice direction, standing order, ethics opinion, statute, regulation, policy, consultation and guidance document on AI in legal practice that we have verified — one row per document… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/legal-ai-regulation-documents.sorrel-T-gemma-3-27b-pt-seed0-documentsdocument-qna-chroma-anyscale-logspython_documentation_codeDataset created using https://github.com/jeffmeloy/py2dataset using the Python code from the following:
https://github.com/ansible/ansible
https://github.com/apache/airflow
https://github.com/arogozhnikov/einops
https://github.com/arviz-devs/arviz
https://github.com/astropy/astropy
https://github.com/biopython/biopython
https://github.com/bjodah/chempy
https://github.com/bokeh/bokehhttps://github.com/CalebBell/thermo
https://github.com/camDavidsonPilon/lifelines
https://github.com/coin-or/pulp… See the full description on the dataset page: https://huggingface.co/datasets/jeffmeloy/python_documentation_code.apex-document-relevance-review
Apex Telecommunications Document Relevance Review (Synthetic)
A fully synthetic eDiscovery / internal-investigation dataset built to test whether an AI agent can
perform first-pass relevance coding on a realistic, adversarial document population.
The scenario
Apex Telecommunications, Inc. suspects Senior Sales Manager Michael Carter of sharing confidential
pricing and customer information with a competitor, Northstar Communications, before and during his
departure… See the full description on the dataset page: https://huggingface.co/datasets/Vrishab80/apex-document-relevance-review.library-documentationThe library documentation retrieval source for code-rag-bench, contains all documentation for Python libraries available on devdocs.io.
sorrel-T2-qwen3-8b-base-seed0-documentspersian-web-document-retrieval
Dataset Summary
Persian Web Document Retrieval is a Persian (Farsi) dataset designed for the Retrieval task. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset consists of real-world queries collected from the Zarrebin search engine and web documents labeled by humans for relevance. It is curated to evaluate model performance in web search scenarios.
Language(s): Persian (Farsi)
Task(s): Retrieval (Web Search)
Source: Collected from Zarrebin… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-web-document-retrieval.sorrel-T2-qwen3-4b-base-seed0-documentsCorrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text.
This dataset is intended to be used to train LLM to perform text-recovery task.
All the errors in text is made solely in Russian sentences, hence ignoring any English sentence.
Texts contains complex formatting, which is common for documents.
Each line contains json object that have array messages value, which consists of role-based conversation.
Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.
