datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs
Liberating 3T of the finest tokens from PDFs
What is this?
As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs.
📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.finepdfs-edu
📚 FinePDFs-Edu
350B+ of highly educational tokens from PDFs 📄
What is it?
📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages.
FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset.
We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.finepdfs-dclm-fineweb-edu-fi
FinePDFs · DCLM · FineWeb-Edu — Finnish (machine-translated)
Finnish machine translation of an English pretraining mixture drawn from
FinePDFs, DCLM, and FineWeb-Edu (documents up to ~900 tokens),
produced as continued-pretraining data for Finnish LLMs.
Translation model: translategemma-27b (Gemma-based 27B translation model)
Documents: ~4,000,000 (80 shards × 50,000)
Language: Finnish (fi)
Format: plain text, one document per row
Originally stored as TFDS-style ArrayRecord… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/finepdfs-dclm-fineweb-edu-fi.finepdfs
Liberating 3T of the finest tokens from PDFs
What is this?
As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs.
📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Compared to HTML… See the full description on the dataset page: https://huggingface.co/datasets/KefranAbg/finepdfs.finepdfs-summaries
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The FinePDFs Summaries subset of MultiSynt is made of summaries LLM-generated with Qwen3-Next-80B-A3B-Instruct for documents from finepdfs.
Work in progress, still generating more data.
The following table shows the data available for each language:
Language
Summaries
Tokens
Disk size
All
838,268,819
247 B
366 GB
deu_Latn
363,671,069
113 B
149 GB
eng_Latn
353,969,370
89 B
162 GB
fra_Latn
27,308… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/finepdfs-summaries.finepdfs-long
HuggingFaceFW/finepdfs long filtered Turkish texts
Source: HuggingFaceFW/finepdfs (config: tur_Latn).
Rows contain 4,000–15,500 characters and passed the
iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft
information-density filters. Selected rows: 140,166.
Generated by process_hf_dataset.py. See summary.json for counts and thresholds.
urdu_finepdfs
What’s inside
data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards.
scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text).
README.md — this file.
If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.finepdfs-et
FinePDFs-et
The Estonian subset of HuggingFaceFW/finepdfs reuploaded for ease of access.
Licensing Information
The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use.
Citation Information
@misc{kydlicek2025finepdfs,
title={FinePDFs},
author={Hynek Kydl{\'\i}{\v{c}}ek and Guilherme Penedo and Leandro von Werra},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/finepdfs-et.FinePDFs-zh
FinePDFs-zh dataset card
Introduction
FinePDFs-zh is a fine-grained classified pdf dataset derived from the cmn_Hani subset of FinePDFs. Each sample is classified into Traditional Chinese, Simplified Chinese, Cantonese, Classical Chinese (Traditional), and Classical Chinese (Simplified) using MusubiAI/ZHLID.
Limitation
Data may be classified under the wrong label due to misclassification by the MusubiAI/ZHLID model. Users are encouraged to perform… See the full description on the dataset page: https://huggingface.co/datasets/MusubiAI/FinePDFs-zh.finepdfs-zhtw
Dataset Card for finepdfs-zhtw
finepdfs-zhtw 是一個由社群協作匯集之台灣繁體中文 PDF 文件集,目前收錄 23 份原始 PDF 檔(~1.1 GB),由 twinkle-ai/fine-pdf-archive 上傳介面收集(此即為本資料集之原始來源 Space)。每筆資料包含 PDF 原始 bytes、貢獻者、PDF 自身之授權、上傳時間戳與雜湊值,作為後續 PDF-to-text pipeline(如 FinePDFs 風格之 OCR / layout parsing)之原始語料源。
Dataset Details
Dataset Description
繁體中文之高品質 PDF 文件(政府公報、教學講義、研究報告、技術手冊等)長期缺乏系統化收集,現有繁中預訓練語料多以 HTML / 純文字為主,對 PDF 內之表格、圖說、版面資訊幾乎無法涵蓋。本資料集受 HuggingFace FinePDFs 之啟發,專注於台灣來源之繁中 PDF,由 Twinkle AI… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/finepdfs-zhtw.
