datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers.js-docspolicy-docs
Public Policy at Hugging Face
AI Policy at Hugging Face is a multidisciplinary and cross-organizational workstream. Instead of being part of a vertical communications or global affairs organization, our policy work is rooted in the expertise of our many researchers and developers, from Ethics and Society Regulars and legal team to machine learning engineers working on healthcare, art, and evaluations.
What we work on is informed by our Hugging Face community needs and experiences… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/policy-docs.docs-imagesdiffusers-images-docsdolma3-6t-sample-100000-docs
dolma3-6t-sample-100000-docs
Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/).
Layout
HCAI-Lab/dolma3-6t-sample-100000-docs/
├── bin_summary.csv
├── sample_contract.json
└── worker_NNNN/
└── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker)
Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.pixmo-docs
PixMo-Docs
We now recommend using CoSyn-400k and CoSyn-point over these
datasets. They are improved versions with more images categories and an improved generation pipeline.
PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.dolma3-6t-sample-5000-docs
dolma3-6t-sample-5000-docs
Materialized stratified sample of 5K docs per bin (2.86M total docs, 5.3B tokens). Seed 42.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_5000_docs
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-5000-docs.dolma3-6t-sample-10000-docs
dolma3-6t-sample-10000-docs
Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_10000_docs
Renamed
2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.protein-docs
Protein Documents (Parquet)
Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata.
Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M
Document Schemes
Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.docslk-news-docsargilla_sdk_docs_raw_unstructured
Dataset info
This dataset contains documentation chunks from repositories (ADD REPOS).
Postprocessing
After some inspection, some chunks contain text too short to be meaningful, so we decided to remove those by removing chunks whose number of tokens (computed
with the same tokenizer of the model to be used for the embeddings) is lower or equal to the 5%:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("BAAI/bge-base-en-v1.5")
df =… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/argilla_sdk_docs_raw_unstructured.the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied.
The filters filtered in this dataset are:
remove_function_no_docstring
remove_class_no_docstring
remove_delete_markers
lk-tourism-weekly-reports-docsdolma3-6t-sample-1000-docs
dolma3-6t-sample-1000-docs
Materialized stratified sample of 1K docs per bin (575K total docs, 1.08B tokens). Seed 42.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_1000_docs
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-1000-docs.cbsl-annual-reports-docslk-dmc-weather-forecasts-docsengineering-docsmarinfold-exp11-protein-docs-seq
marinfold-exp11-pdocs-seq
Sequence-only derivative of
eczech/marinfold-exp11-protein-docs.
For every row, the document field has been reduced to just the amino-acid sequence
portion: the <begin_sequence> tag followed by the per-residue three-letter tokens
(e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1>
document-type prefix and everything from <begin_statements> onward (contacts and
distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.marinfold-exp11-protein-docs
marinfold-exp11-pdocs
Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from
timodonnell/protein-docs,
partitioned by the source round column:
Config
Source rounds
Approx rows
high
round 0
~1.68M
medium
round 1
~1.42M
low
round 2–4
~2.29M
Train/val/test split assignment is inherited from the source dataset (leakage-resistant
structural-cluster hashing). All columns from the source are preserved; rows are simply
partitioned by round.
See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
lk-tourism-monthly-reports-docsdolma3-6t-sample-500-docs
dolma3-6t-sample-500-docs
Materialized stratified sample of 500 docs per bin (288K total docs, 539M tokens) drawn from the deduplicated 6T Dolma3 corpus. Seed 42. Matching companion bucket at hf://buckets/HCAI-Lab/dolma3-6t-sample-500-docs.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-500-docs.n8n-docs-v2
n8n Docs
This repository hosts the documentation for n8n, an extendable workflow automation tool which enables you to connect anything to everything. The documentation is live at docs.n8n.io.
Previewing and building the documentation locally
Prerequisites
Python 3.8 or above
Pip
n8n recommends using a virtual environment when working with Python, such as venv.
Follow the recommended configuration and auto-complete guidance for the theme. This will help when… See the full description on the dataset page: https://huggingface.co/datasets/Jeice/n8n-docs-v2.DocSynth300K
DocSynth300K is a large-scale and diverse document layout analysis pre-training dataset, which can largely boost model performance.
Data Download
Use following command to download dataset(about 113G):
from huggingface_hub import snapshot_download
# Download DocSynth300K
snapshot_download(repo_id="juliozhao/DocSynth300K", local_dir="./docsynth300k-hf", repo_type="dataset")
# If the download was disrupted and the file is not complete, you can resume the download… See the full description on the dataset page: https://huggingface.co/datasets/juliozhao/DocSynth300K.lamini_docs
Dataset Card for "lamini_docs"
More Information needed
math-docs-dataset
Mathematical Documents Dataset
This dataset contains 36,661 scientific documents with OCR-extracted text and mathematical content probability scores.
Documents were filtered from the CommonCrawl PDF corpus based on mathematical content probability.
Quick Start
from datasets import load_dataset
import json
# Load metadata
with open("metadata.jsonl") as f:
for line in f:
doc = json.loads(line)
doc_id = doc['doc_id']
# Read extracted… See the full description on the dataset page: https://huggingface.co/datasets/Monketoo/math-docs-dataset.DocStruct4Mlk-dmc-river-water-level-and-flood-warnings-docssfm-midtraining-blocklist-filtered-docs-20251123-0747
