CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceTB /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.tabular100M<n<1B486 likes54k downloads2y agoHugging Face02Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face03chengjunyan1 /smollm-12.5-corpus SmolLM-1/8-Corpus Around 1/8 upper-quality subset of SmolLM Corpus for training Chinchilla-optimal GPT-2 scale (sub 1.5B) models, which is a good scale for verifying a model architecture under the scaling laws. Firstly filtered samples with int_score >=4 from FineWeb-edu-dedup, then keep the training mixture with the same distribution from SmolLM. In which FineWeb-Edu-dedup occupies around 70% of the corpus. Then sample other dataset based on the mixture ratios respectively. For… See the full description on the dataset page: https://huggingface.co/datasets/chengjunyan1/smollm-12.5-corpus.tabular100M<n<1B7 likes1.8k downloads2y agoHugging Face04Trelis /smollm-corpus-2percenttext1M<n<10M0 likes483 downloads2y agoHugging Face05styal /smollm-corpus-3.5M A very small version of smollm-corpus (+finemath-4plus) for experimenting llm pre-training. cosmopedia-v2 (1M rows) fineweb-edu-dedup (1M rows) python-edu (0.5M rows) finemath-4plus (1M rows) tabular1M<n<10M2 likes452 downloads2y agoHugging Face06enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes340 downloads8mo agoHugging Face07enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes259 downloads8mo agoHugging Face08oieieio /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/oieieio/smollm-corpus.tabular100M<n<1B0 likes226 downloads1y agoHugging Face09david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes170 downloads4mo agoHugging Face10Mgmgrand420 /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/smollm-corpus.tabular100M<n<1B0 likes77 downloads8mo agoHugging Face11styal /very-smollm-corpus-0.5Mtabular100K<n<1M0 likes69 downloads2y agoHugging Face12ecreeth /1b-smollm-corpus SmolLM-Corpus — 1B Token Subset A curated 1-billion-token English pretraining corpus sampled from HuggingFaceTB/smollm-corpus, designed for training small language models (~20M parameters). Dataset Composition Source Ratio Tokens Documents FineWeb-Edu (dedup) 87% ~870M 849,577 Cosmopedia v2 13% ~130M 161,889 Total 100% ~1B 1,011,466 Rationale for the Split The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page: https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.texttext-generation1M<n<10M1 likes60 downloads4mo agoHugging Face13hhoenjet /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/hhoenjet/smollm-corpus.tabular100M<n<1B0 likes56 downloads7mo agoHugging Face14Jakemu /smollm-corpus-shuffletext1M<n<10M0 likes46 downloads8mo agoHugging Face15Jakemu /smollm-corpus-and-FineWeb2-ko-synth-shuffle-20btext10M<n<100M0 likes46 downloads7mo agoHugging Face16david-thrower /HelixLM-s512-instruct-smollm-stock-corpus-1-7m-token david-thrower/HelixLM-s512-instruct-smollm-stock-corpus-1-7m-token Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/HelixLM-s512-instruct-smollm-stock-corpus-1-7m-token.text1K<n<10K0 likes17 downloads4mo agoHugging Face17Obscure-Entropy /smollm-corpus-hutext100K<n<1M0 likes14 downloads1y agoHugging Face18eliebak /very-smollm-corpustext1M<n<10M3 likes13 downloads2y agoHugging Face19BEE-spoke-data /smollm-corpus-pythongated smollm-corpus - python A version of the python-edu subset with the text added tabulartext-generation10M<n<100M0 likes6 downloads9mo agoHugging Face20jsanzolac /smollm_corpustext1M<n<10M0 likes4 downloads5mo agoHugging Face21Jakemu /smollm-corpus-and-FineWeb2-ko-synth-shuffletext1M<n<10M0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.