CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /finepdfs Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.tabulartext-generation100M<n<1B942 likes43k downloads6mo agoHugging Face02HuggingFaceFW /finepdfs-edu 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.tabulartext-generation10M<n<100M98 likes17k downloads11mo agoHugging Face03KefranAbg /finepdfs Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to HTML… See the full description on the dataset page: https://huggingface.co/datasets/KefranAbg/finepdfs.tabulartext-generation100M<n<1B1 likes160 downloads10mo agoHugging Face04Ba2han /finepdfs-long HuggingFaceFW/finepdfs long filtered Turkish texts Source: HuggingFaceFW/finepdfs (config: tur_Latn). Rows contain 4,000–15,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 140,166. Generated by process_hf_dataset.py. See summary.json for counts and thresholds. tabulartext-generation100K<n<1M0 likes47 downloads2mo agoHugging Face05humair025 /urdu_finepdfs What’s inside data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards. scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text). README.md — this file. If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.tabulartext-classification100K<n<1M0 likes28 downloads10mo agoHugging Face06tartuNLP /finepdfs-et FinePDFs-et The Estonian subset of HuggingFaceFW/finepdfs reuploaded for ease of access. Licensing Information The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use. Citation Information @misc{kydlicek2025finepdfs, title={FinePDFs}, author={Hynek Kydl{\'\i}{\v{c}}ek and Guilherme Penedo and Leandro von Werra}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/finepdfs-et.tabulartext-generation100K<n<1M0 likes26 downloads9mo agoHugging Face07MusubiAI /FinePDFs-zh FinePDFs-zh dataset card Introduction FinePDFs-zh is a fine-grained classified pdf dataset derived from the cmn_Hani subset of FinePDFs. Each sample is classified into Traditional Chinese, Simplified Chinese, Cantonese, Classical Chinese (Traditional), and Classical Chinese (Simplified) using MusubiAI/ZHLID. Limitation Data may be classified under the wrong label due to misclassification by the MusubiAI/ZHLID model. Users are encouraged to perform… See the full description on the dataset page: https://huggingface.co/datasets/MusubiAI/FinePDFs-zh.tabulartext-generation1M<n<10M4 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.