CoolFace
19 results

finepdfs

HuggingFaceFW /finepdfs Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.tabulartext-generation100M<n<1B942 likes43k downloads6mo agoHugging FaceHuggingFaceFW /finepdfs_lang_classificationtabular1M<n<10M4 likes19k downloads11mo agoHugging FaceHuggingFaceFW /finepdfs-edu 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.tabulartext-generation10M<n<100M98 likes17k downloads11mo agoHugging FaceHuggingFaceFW /finepdfs_fw_edu_labeledtext10M<n<100M6 likes4.4k downloads1y agoHugging FaceHuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs 100BT FinePDFs ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT FineWeb-Edu ~20B The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M2 likes2.9k downloads7mo agoHugging FaceHuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs-Edu 100BT FinePDFs-Edu ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M0 likes2.5k downloads7mo agoHugging Face