CoolFace
Datasetpublic

Mgmgrand420/smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/smollm-corpus.

sourceHugging Faceodc-byupdated 8mo agoView on Hugging Face
0likes82downloads
settings

This repository belongs to Mgmgrand420 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesmollm-corpus
visibilitypublic
licenceodc-by
gatedno
ownerMgmgrand420
Account settings
Mgmgrand420/smollm-corpus · CoolFace