CoolFace
Datasetpublic

melaniaghirda/fineweb-edu-subset

Small slice of the original HuggingFaceFW/fineweb-edu, data/CC-MAIN-2025-26. Files: first 10 .parquet files, split=train, columns="text", applied filters: "language_score">=0.9. Pipeline: Dataset will be further tokenized and used to train a 124M GPT model.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes519downloads
Dataset Card

Small slice of the original HuggingFaceFW/fineweb-edu, data/CC-MAIN-2025-26.<br> Files: first 10 .parquet files, split=train, columns="text", applied filters: "language_score">=0.9.<br> Pipeline: Dataset will be further tokenized and used to train a 124M GPT model.