datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.MINT-1T-PDF-CC-2023-40
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-40.MINT-1T-PDF-CC-2023-06
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-06.MINT-1T-PDF-CC-2024-18
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-18.MINT-1T-PDF-CC-2023-14
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.SCPWiki-Cleaned-PDF-ArchivesMINT-1T-PDF-CC-2023-50
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.biorXiv-pdf
BiorXiv Pdf
BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets.
BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.llm-jp-corpus-v4-ja_warp_pdf
llm-jp-corpus-v4 — ja_warp_pdf
Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_warp_pdf
Files: 513 × jsonl.gz (73.8 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.ocr-pdf-degraded
OCR-PDF-Degraded Dataset
Overview
This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments.
Purpose
Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.Pashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.StenCore-PDF
StenCore — FinePDFs-Edu Curated
By StentorLabs
StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting.
⚠️ Privacy &… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.Books-pdfnapierone-pdf-nanonets-s
NapierOne PDFs: OCR'd by nanonets-s
PDFs from NapierOne (see 'pdf-total' in the napierone aws bucket) converted to text with nanonets-s using this code
contains results for all 4978 unique PDFs
raw config is unmodified from model output, the default config has been post-processed with mdformat
Citation
@article{DAVIES2022301330,
title = {NapierOne: A modern mixed file data set alternative to Govdocs1},
journal = {Forensic Science International: Digital… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/napierone-pdf-nanonets-s.curatorkit-testrun-PDF
curatorkit-testrun-PDF
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 06:27 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca")
napierone-pdf-olmOCR
NapierOne PDFs - converted with olmOCR
PDFs from NapierOne (see 'pdf-total' in the napierone aws bucket) converted to text via the olmOCR pipeline
outputs with <= 100 chars and 'low quality pdfs' were filtered out
the 4228 rows in default config represent approx 93,595 input PDF pages
Citation
@article{DAVIES2022301330,
title = {NapierOne: A modern mixed file data set alternative to Govdocs1},
journal = {Forensic Science International: Digital Investigation}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/napierone-pdf-olmOCR.khmer-education-pdf-cleaned
Khmer Education PDF - Role-Playing Teaching Methods (Cleaned)
Dataset Description
This dataset contains cleaned and Unicode-corrected text extracted from a 48-page Khmer educational PDF about role-playing teaching methods (ការបងៀនតាមវិធីសម្មែងតួ). The text has undergone comprehensive Unicode repair to fix 537 orphaned COENG characters and other extraction artifacts.
Dataset Summary
Language: Khmer (km)
Source: Educational PDF - "New Generation Pedagogical… See the full description on the dataset page: https://huggingface.co/datasets/khopilot/khmer-education-pdf-cleaned.pdfsys-page-v2-demo
pdfsys.page/v2 — 格式演示数据集
pdfsys.page/v2 是 pdfsystem_mnbvc
的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。
这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、
三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。
来源提示:这里的 PDF 页来自 OmniDocBench
与 olmOCR-bench 两个公开
benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些
文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。
一句话设计
一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的;
页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强;
图像像素要么是裁剪图、要么是整页光栅,二选一。
里面有什么
config
行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.napierone-pdf-raw
BEE-spoke-data/napierone-pdf-raw
NapierOne PDF files converted with marker.
detected languages
Counter({'en': 4665,
'nl': 2,
'fi': 7,
'fr': 8,
'cy': 54,
'sq': 1,
'it': 1,
'unknown-error': 5,
'sk': 1,
'es': 2,
'de': 3,
'ro': 1,
'pl': 1,
'zh': 1,
'so': 1,
'ml': 1})
nist-coa-pdfThis is a set of chemical levels measured in NIST SRMs as reported in their Certificates of Analysis (COA) PDF documents. Data was manually extracted by Dr. Lane C. Sander of NIST.codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples
FIneMix Dataset:
This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m.
It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources.
Data Sources and selection:
50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.edrXiv-pdfEdArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the Education domain. Managed by the Centre of Open Science and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in Education.
As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in traditional science but also across a broad range of scientific disciplines.… See the full description on the dataset page: https://huggingface.co/datasets/laion/edrXiv-pdf.PsyArXiv-pdfPsyArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the psychiatry domain. Managed by the Centre of Open Science and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in Education.
As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in traditional science but also across a broad range of scientific disciplines.… See the full description on the dataset page: https://huggingface.co/datasets/laion/PsyArXiv-pdf.metaArXiv-pdfEdArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the Education domain. Managed by the Berkeley Initiative for
Transparency in the Social Sciences and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in research transparency and reproducibility.
As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in… See the full description on the dataset page: https://huggingface.co/datasets/laion/metaArXiv-pdf.
