CoolFace
Datasetpublic

HuggingFaceFW/finepdfs

Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
942likes41kdownloads
settings

This repository belongs to HuggingFaceFW on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefinepdfs
visibilitypublic
licenceodc-by
gatedno
ownerHuggingFaceFW
Account settings
HuggingFaceFW/finepdfs · CoolFace