CoolFace
Datasetpublic

KefranAbg/finepdfs

Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/KefranAbg/finepdfs.

sourceHugging Faceodc-byupdated 10mo agoView on Hugging Face
1likes116downloads
../
file000_00000.parquet19 KBdownload

KefranAbg/finepdfs · main · files are served by the source, never re-hosted here