datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs
Liberating 3T of the finest tokens from PDFs
What is this?
As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs.
📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.finepdfs-edu
📚 FinePDFs-Edu
350B+ of highly educational tokens from PDFs 📄
What is it?
📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages.
FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset.
We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.finepdfs
Liberating 3T of the finest tokens from PDFs
What is this?
As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs.
📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Compared to HTML… See the full description on the dataset page: https://huggingface.co/datasets/KefranAbg/finepdfs.finepdfs-long
HuggingFaceFW/finepdfs long filtered Turkish texts
Source: HuggingFaceFW/finepdfs (config: tur_Latn).
Rows contain 4,000–15,500 characters and passed the
iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft
information-density filters. Selected rows: 140,166.
Generated by process_hf_dataset.py. See summary.json for counts and thresholds.
urdu_finepdfs
What’s inside
data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards.
scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text).
README.md — this file.
If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.finepdfs-et
FinePDFs-et
The Estonian subset of HuggingFaceFW/finepdfs reuploaded for ease of access.
Licensing Information
The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use.
Citation Information
@misc{kydlicek2025finepdfs,
title={FinePDFs},
author={Hynek Kydl{\'\i}{\v{c}}ek and Guilherme Penedo and Leandro von Werra},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/finepdfs-et.FinePDFs-zh
FinePDFs-zh dataset card
Introduction
FinePDFs-zh is a fine-grained classified pdf dataset derived from the cmn_Hani subset of FinePDFs. Each sample is classified into Traditional Chinese, Simplified Chinese, Cantonese, Classical Chinese (Traditional), and Classical Chinese (Simplified) using MusubiAI/ZHLID.
Limitation
Data may be classified under the wrong label due to misclassification by the MusubiAI/ZHLID model. Users are encouraged to perform… See the full description on the dataset page: https://huggingface.co/datasets/MusubiAI/FinePDFs-zh.
