datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.MMLongBench-docprotein-docs
Protein Documents (Parquet)
Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata.
Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M
Document Schemes
Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled.
dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval.
The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied.
The filters filtered in this dataset are:
remove_function_no_docstring
remove_class_no_docstring
remove_delete_markers
docmath-eval-failures-200
DocMath-Eval Failures 200: Agent Benchmark & Leaderboard
A curated benchmark of 200 challenging financial math questions that leading AI models
failed to answer correctly, with comprehensive evaluation results from multiple AI agents.
Leaderboard
Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring.
Rank
Agent
Model
Exact Match
Judge: Exact
Judge: Approx
Judge: Total
Wrong
Avg Duration
Avg Tool Calls
1
TRAE Agent
Opus 4.5
98/200 (49.0%)
96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.Agilex_Cobot_Magic_zip_up_the_document_bag
Agilex_Cobot_Magic_zip_up_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pull
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.cc100-documents
cc100-documents
This dataset is a restructured version of the CC-100 (statmt/cc100) dataset.
In the original dataset, each instance corresponds to a single paragraph (or a document boundary).
In this version, the data has been reformed so that each instance corresponds to a single, complete document.
This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library.
Languages
The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled.
marinfold-exp11-protein-docs-seq
marinfold-exp11-pdocs-seq
Sequence-only derivative of
eczech/marinfold-exp11-protein-docs.
For every row, the document field has been reduced to just the amino-acid sequence
portion: the <begin_sequence> tag followed by the per-residue three-letter tokens
(e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1>
document-type prefix and everything from <begin_statements> onward (contacts and
distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.RMC-AIDA-L_organise_the_document_bag
RMC-AIDA-L_organise_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
pull
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.XL-DocBench
XL-DocBench
Evidence-grounded reasoning across hundreds or thousands of pages.
Fully verified by 194 human experts.
Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡,
Bei Liu2,*, Yifan Yang2, Qi Dai2,
Ruichun Ma2, Kai Qiu2, Yunsheng Li2,
Dongdong Chen2, Chong Luo2,
Zhenzhong Chen1, Baining Guo2
1Wuhan University 2Microsoft
†Equal contribution ‡Work done during an internship at MSRA
*Project leader
Project Page ·
Paper ·
Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.marinfold-exp11-protein-docs
marinfold-exp11-pdocs
Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from
timodonnell/protein-docs,
partitioned by the source round column:
Config
Source rounds
Approx rows
high
round 0
~1.68M
medium
round 1
~1.42M
low
round 2–4
~2.29M
Train/val/test split assignment is inherited from the source dataset (leakage-resistant
structural-cluster hashing). All columns from the source are preserved; rows are simply
partitioned by round.
See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.r2e-dockers-rllm-v1wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.Chinese_Debate_Documents
Dataset Card for Chinese Debate Documents
ASR-transcribed corpus of competitive Mandarin Chinese university debates,
speaker-segmented and timestamped, with topic / round / team metadata parsed
from the source filenames.
Loading the Dataset
from datasets import load_dataset
ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train")
print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"])
for seg in ds[0]["segments"][:3]:
print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.dockersv1the-stack-dedup-python-filtered-docstrings-gpt2text-code-galeras-code-generation-from-docstring-3k-dedupedUDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.data_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.r2e-dockers-v3r2e-dockers-v2conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.finemed-fr
FineMed-fr
🤗 Blog |
📄 Paper |
💻 Code |
🌐 FineMed |
🩺 DoctoBERT
📚 Introduction
FineMed-fr is a large, openly available corpus of French medical text for language-model pretraining: 21.1M documents and 19.2B words of real-world medical writing, annotated along several quality axes.
The corpus is drawn from three heterogeneous open-web sources (FineWeb-2,
FinePDFs, and
FineWiki), which together provide the scale, source
diversity, and stylistic range… See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-fr.docx-corpus
docx-corpus
The largest classified corpus of Word documents. 736K+ .docx files from the public web, classified into 10 document types and 9 topics across 76 languages.
Dataset Description
This dataset contains metadata for publicly available .docx files collected from the web. Each document has been classified by document type and topic using a two-stage pipeline: LLM labeling (Claude) of a stratified sample, followed by fine-tuned XLM-RoBERTa classifiers applied at… See the full description on the dataset page: https://huggingface.co/datasets/superdoc-dev/docx-corpus.gmane_dataset-docus-court-opinions-dockets-judges-dataset
US Court Opinions Metadata, Dockets & Judges (CourtListener)
10M opinion clusters, 70M dockets and 16K judges from official CourtListener / Free Law Project bulk data as metadata + derived-signals tables — citation graph, company litigation profiles; no opinion full text.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-court-opinions-dockets-judges-dataset… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-court-opinions-dockets-judges-dataset.doc2
doc2lora_full_p0p1_text_v1
Text-format parquet dataset for document-to-LoRA (doc2lora) continual-learning pretraining.
Plain-text fields (context/prompt/response/suffix) plus Qwen3.5-4B token id lists
(*_tokens_qwen35) for arbitrary-model re-tokenization.
Train split: train/ — 3287 parquet shards, 32,703,729 rows, ~7.7 GB.
reconstruct_conversation: prompt + response + suffix ; supervised_text: response.
Auto-generated card. Uploaded via hf-mirror.com.
