CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /finepdfs Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.tabulartext-generation100M<n<1B943 likes44k downloads6mo agoHugging Face02HuggingFaceFW /finepdfs-edu 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.tabulartext-generation10M<n<100M98 likes16k downloads11mo agoHugging Face03Finnish-NLP /finepdfs-dclm-fineweb-edu-fi FinePDFs · DCLM · FineWeb-Edu — Finnish (machine-translated) Finnish machine translation of an English pretraining mixture drawn from FinePDFs, DCLM, and FineWeb-Edu (documents up to ~900 tokens), produced as continued-pretraining data for Finnish LLMs. Translation model: translategemma-27b (Gemma-based 27B translation model) Documents: ~4,000,000 (80 shards × 50,000) Language: Finnish (fi) Format: plain text, one document per row Originally stored as TFDS-style ArrayRecord… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/finepdfs-dclm-fineweb-edu-fi.texttext-generation1M<n<10M1 likes286 downloads2mo agoHugging Face04KefranAbg /finepdfs Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to HTML… See the full description on the dataset page: https://huggingface.co/datasets/KefranAbg/finepdfs.tabulartext-generation100M<n<1B1 likes116 downloads10mo agoHugging Face05MultiSynt /finepdfs-summaries MultiSynt MultiSynt is an open multilingual synthetic dataset. The FinePDFs Summaries subset of MultiSynt is made of summaries LLM-generated with Qwen3-Next-80B-A3B-Instruct for documents from finepdfs. Work in progress, still generating more data. The following table shows the data available for each language: Language Summaries Tokens Disk size All 838,268,819 247 B 366 GB deu_Latn 363,671,069 113 B 149 GB eng_Latn 353,969,370 89 B 162 GB fra_Latn 27,308… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/finepdfs-summaries.texttext-generation1B<n<10B2 likes70 downloads7mo agoHugging Face06Ba2han /finepdfs-long HuggingFaceFW/finepdfs long filtered Turkish texts Source: HuggingFaceFW/finepdfs (config: tur_Latn). Rows contain 4,000–15,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 140,166. Generated by process_hf_dataset.py. See summary.json for counts and thresholds. tabulartext-generation100K<n<1M0 likes40 downloads2mo agoHugging Face07humair025 /urdu_finepdfs What’s inside data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards. scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text). README.md — this file. If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.tabulartext-classification100K<n<1M0 likes29 downloads10mo agoHugging Face08tartuNLP /finepdfs-et FinePDFs-et The Estonian subset of HuggingFaceFW/finepdfs reuploaded for ease of access. Licensing Information The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use. Citation Information @misc{kydlicek2025finepdfs, title={FinePDFs}, author={Hynek Kydl{\'\i}{\v{c}}ek and Guilherme Penedo and Leandro von Werra}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/finepdfs-et.tabulartext-generation100K<n<1M0 likes27 downloads9mo agoHugging Face09MusubiAI /FinePDFs-zh FinePDFs-zh dataset card Introduction FinePDFs-zh is a fine-grained classified pdf dataset derived from the cmn_Hani subset of FinePDFs. Each sample is classified into Traditional Chinese, Simplified Chinese, Cantonese, Classical Chinese (Traditional), and Classical Chinese (Simplified) using MusubiAI/ZHLID. Limitation Data may be classified under the wrong label due to misclassification by the MusubiAI/ZHLID model. Users are encouraged to perform… See the full description on the dataset page: https://huggingface.co/datasets/MusubiAI/FinePDFs-zh.tabulartext-generation1M<n<10M4 likes12 downloads1y agoHugging Face10lianghsun /finepdfs-zhtwgated Dataset Card for finepdfs-zhtw finepdfs-zhtw 是一個由社群協作匯集之台灣繁體中文 PDF 文件集,目前收錄 23 份原始 PDF 檔(~1.1 GB),由 twinkle-ai/fine-pdf-archive 上傳介面收集(此即為本資料集之原始來源 Space)。每筆資料包含 PDF 原始 bytes、貢獻者、PDF 自身之授權、上傳時間戳與雜湊值,作為後續 PDF-to-text pipeline(如 FinePDFs 風格之 OCR / layout parsing)之原始語料源。 Dataset Details Dataset Description 繁體中文之高品質 PDF 文件(政府公報、教學講義、研究報告、技術手冊等)長期缺乏系統化收集,現有繁中預訓練語料多以 HTML / 純文字為主,對 PDF 內之表格、圖說、版面資訊幾乎無法涵蓋。本資料集受 HuggingFace FinePDFs 之啟發,專注於台灣來源之繁中 PDF,由 Twinkle AI… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/finepdfs-zhtw.texttext-generationn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.