datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docmath-eval-failures-200
DocMath-Eval Failures 200: Agent Benchmark & Leaderboard
A curated benchmark of 200 challenging financial math questions that leading AI models
failed to answer correctly, with comprehensive evaluation results from multiple AI agents.
Leaderboard
Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring.
Rank
Agent
Model
Exact Match
Judge: Exact
Judge: Approx
Judge: Total
Wrong
Avg Duration
Avg Tool Calls
1
TRAE Agent
Opus 4.5
98/200 (49.0%)
96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.protein-docs
Protein Documents (Parquet)
Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata.
Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M
Document Schemes
Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval.
The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.cc100-documents
cc100-documents
This dataset is a restructured version of the CC-100 (statmt/cc100) dataset.
In the original dataset, each instance corresponds to a single paragraph (or a document boundary).
In this version, the data has been reformed so that each instance corresponds to a single, complete document.
This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library.
Languages
The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.marinfold-exp11-protein-docs-seq
marinfold-exp11-pdocs-seq
Sequence-only derivative of
eczech/marinfold-exp11-protein-docs.
For every row, the document field has been reduced to just the amino-acid sequence
portion: the <begin_sequence> tag followed by the per-residue three-letter tokens
(e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1>
document-type prefix and everything from <begin_statements> onward (contacts and
distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.Chinese_Debate_Documents
Dataset Card for Chinese Debate Documents
ASR-transcribed corpus of competitive Mandarin Chinese university debates,
speaker-segmented and timestamped, with topic / round / team metadata parsed
from the source filenames.
Loading the Dataset
from datasets import load_dataset
ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train")
print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"])
for seg in ds[0]["segments"][:3]:
print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.UDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.marinfold-exp11-protein-docs
marinfold-exp11-pdocs
Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from
timodonnell/protein-docs,
partitioned by the source round column:
Config
Source rounds
Approx rows
high
round 0
~1.68M
medium
round 1
~1.42M
low
round 2–4
~2.29M
Train/val/test split assignment is inherited from the source dataset (leakage-resistant
structural-cluster hashing). All columns from the source are preserved; rows are simply
partitioned by round.
See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.finemed-fr
FineMed-fr
🤗 Blog |
📄 Paper |
💻 Code |
🌐 FineMed |
🩺 DoctoBERT
📚 Introduction
FineMed-fr is a large, openly available corpus of French medical text for language-model pretraining: 21.1M documents and 19.2B words of real-world medical writing, annotated along several quality axes.
The corpus is drawn from three heterogeneous open-web sources (FineWeb-2,
FinePDFs, and
FineWiki), which together provide the scale, source
diversity, and stylistic range… See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-fr.conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.finemed-rephrased-fr
FineMed-rephrased-fr
🤗 Blog |
📄 Paper |
💻 Code |
🌐 FineMed |
🩺 DoctoBERT
📚 Introduction
FineMed-rephrased-fr is a signal-amplifying rephrasing of FineMed-fr: 13.6M documents and 4.5B words of LLM-rephrased French medical text. An LLM rewrites each source document into a faithful variant that raises medical-term density and broadens the co-occurrence context around each medical concept, using an adapted Massive Genre-Audience (MGA) reformulation.
As… See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-rephrased-fr.document-corpus-v3-open
Document Corpus v3 Open
document-corpus-v3-open is the redistribution-compatible slice of the exact
byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT
experiments. It contains 869,739 filtered documents and
2.192 GB of UTF-8 text before Parquet compression.
This is not the complete internal document-corpus-v3. Restricted,
unknown-license, and share-alike sources were excluded conservatively. Every
included row comes from an upstream dataset whose card… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.bureau-docket
The Bureau Corpus — Hyperstition Filings & Records (BIS-F-2026)
This is a work of fiction — a designed art object. Every entry is invented. Nothing here is a prediction, forecast, offer, or advice. The Bureau of Imaginary Solutions does not exist. First filed at ubu.numetal.xyz; companion model at CLINAMEN-45B-A9B.
The public records of the Bureau of Imaginary Solutions — 1,130 documents: 965 hyperstition filings plus 165 Bureau records (memos, appeals, syzygy essays, a… See the full description on the dataset page: https://huggingface.co/datasets/gokhanturhan/bureau-docket.bioreason-pro-sft-reasoning-documents
BioReason-Pro SFT Reasoning Documents
Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training.
The upstream dataset wanglab/bioreason-pro-sft-reasoning-data
ships the assistant side of each training example (reasoning, final_answer) alongside the raw
biological context columns, but not the assembled prompt. The prompt cannot be recovered from the
data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.common-crawl-docx-sample
Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
Source
Common Crawl release: CC-MAIN-YYYY-NN
Source index: Common Crawl URL Index
Pipeline: marin-community/marin
Pipeline revision: REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type,
or a .docx URL suffix. Only successful, non-truncated index records were
eligible.
Processing
The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.rust-cli-docs-corpus
Rust CLI Documentation Corpus
A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools.
Dataset Description
This corpus follows the Toyota Way principles and Popperian falsification methodology.
Statistics
Total entries: 80
Source repositories: 0
Validation score: 96/100
Supported Tasks
Documentation Generation: Generate Rust doc comments from code signatures
Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead.
文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。
推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量
经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式
chat格式为chatml
whylab-gemini-2-5-docker-validation
🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced.
DOI
This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/abby2231/indian-legal-documents.swe_doc_gen_locate_swebench_1500
SWE-Doc-Gen-Locate Dataset (SWE-Bench 1500 entries)
A dataset for evaluating an agent's ability to locate a target Python function/class based on its functionality description and add a docstring.
Task Description
Given: A Python repository and a description of a function/class (NO name, NO file path)
Agent must:
Search the codebase to find where the target function/class is defined
Read the implementation to understand its behavior
Generate and add an appropriate… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/swe_doc_gen_locate_swebench_1500.swe_doc_gen_locate_2000
SWE-Doc-Gen-Locate Dataset (2000 entries)
A dataset for evaluating an agent's ability to locate a target Python function/class based on its functionality description and add a docstring.
Task Description
Given: A Python repository and a description of a function/class (NO name, NO file path)
Agent must:
Search the codebase to find where the target function/class is defined
Read the implementation to understand its behavior
Generate and add an appropriate docstring… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/swe_doc_gen_locate_2000.nyan_documents
Nyan documents
Documents scraped for НЯН Telegram channel from March 2022 to December 2023. The dataset includes documents from 100+ different Telegram news channels.
Usage
pip3 install datasets
from datasets import load_dataset
for row in load_dataset("NyanNyanovich/nyan_documents", split="train", streaming=True):
print(row)
break
Other datasets
Documents (this dataset): https://huggingface.co/datasets/NyanNyanovich/nyan_documents
Clusters:… See the full description on the dataset page: https://huggingface.co/datasets/NyanNyanovich/nyan_documents.swe_doc_gen_locate_1500
SWE-Doc-Gen-Locate Dataset (1500 entries)
A dataset for evaluating an agent's ability to locate a target Python function/class based on its functionality description and add a docstring.
Task Description
Given: A Python repository and a description of a function/class (NO name, NO file path)
Agent must:
Search the codebase to find where the target function/class is defined
Read the implementation to understand its behavior
Generate and add an appropriate docstring… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/swe_doc_gen_locate_1500.docvqa-nanochat
DocVQA for Nanochat
Single-page document QA dataset processed for nanochat fine-tuning.
Description
This dataset is derived from pixparse/docvqa-single-page-questions and has been processed for efficient fine-tuning of small language models with limited context windows.
Modifications from Source
OCR truncation: Answer-priority truncation ensures the answer is always present in the truncated context. Lines containing the answer are prioritized, then surrounding… See the full description on the dataset page: https://huggingface.co/datasets/morgan/docvqa-nanochat.Bangladeshi_Doctor_List
Bangladeshi_Doctor_List Dataset
This Dataset contains all verified and authorized Docto information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact
mahadise01@gmail.com
Linkdin:… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_Doctor_List.docentesDC-curado
docentesDC Curado
Descrição
Este dataset é uma versão curada e consolidada do conjunto de dados vickminari/docentesDC.
O conjunto contém textos associados a professores do Departamento de Computação. A versão curada foi preparada para experimentos acadêmicos envolvendo:
geração de texto;
ajuste fino de modelos de linguagem;
recuperação de informação;
sistemas RAG;
avaliação de respostas baseadas em documentos.
O processamento priorizou a preservação do conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/wesleycoutinhodev/docentesDC-curado.ml-design-doc-reviewer-data
ml-system-design/ml-design-doc-reviewer-data (v1.0.0)
Evaluation artifacts for the ML Design Doc Reviewer project.
Layout
Path
Description
manifest/sample_manifest.csv
Stratified 100-case sample manifest
manifest/error_topology.csv
Controlled error taxonomy for flawed docs
raw/
Raw markdown exports, metadata sidecars, OCR image blocks
raw/images/
Downloaded article images
normalized/
Canonical 14-section ML design documents
flawed/
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/ml-system-design/ml-design-doc-reviewer-data.
