CoolFace
Datasetpublic

Web3Survivor/Survivor

๐Ÿ“š FinePDFs-Edu 350B+ of highly educational tokens from PDFs ๐Ÿ“„ What is it? ๐Ÿ“š FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from ๐Ÿ“„ FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retainโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.

sourceHugging Faceodc-byupdated 10mo agoView on Hugging Face
2likes383downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone elseโ€™s repository from here would need an authorised integration and the account holderโ€™s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Web3Survivor/Survivor ยท CoolFace