datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.cosmopedia-v2-4B
Dataset: cosmopedia-v2-4B
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-4B/stage_1/tmp.
cosmopedia-v2-noeod_10b
Dataset: cosmopedia-v2-noeod_10b
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_10b/stage_1.
