melaniaghirda/fineweb-edu-subset
Small slice of the original HuggingFaceFW/fineweb-edu, data/CC-MAIN-2025-26. Files: first 10 .parquet files, split=train, columns="text", applied filters: "language_score">=0.9. Pipeline: Dataset will be further tokenized and used to train a 124M GPT model.
0563
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face