CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceTB /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.tabular100M<n<1B486 likes54k downloads2y agoHugging Face02Avelina /smollm-corpus SmolLM-Corpus: Now shuffled and sharded! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23 containing 1000 shards… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus.text-generation100M<n<1B5 likes21k downloads2y agoHugging Face03Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face04chengjunyan1 /smollm-12.5-corpus SmolLM-1/8-Corpus Around 1/8 upper-quality subset of SmolLM Corpus for training Chinchilla-optimal GPT-2 scale (sub 1.5B) models, which is a good scale for verifying a model architecture under the scaling laws. Firstly filtered samples with int_score >=4 from FineWeb-edu-dedup, then keep the training mixture with the same distribution from SmolLM. In which FineWeb-Edu-dedup occupies around 70% of the corpus. Then sample other dataset based on the mixture ratios respectively. For… See the full description on the dataset page: https://huggingface.co/datasets/chengjunyan1/smollm-12.5-corpus.tabular100M<n<1B7 likes1.8k downloads2y agoHugging Face05OpenMOSS-Team /MHA2MLA-corpus-smollm0 likes1.3k downloads1y agoHugging Face06Trelis /smollm-corpus-2percenttext1M<n<10M0 likes483 downloads2y agoHugging Face07styal /smollm-corpus-3.5M A very small version of smollm-corpus (+finemath-4plus) for experimenting llm pre-training. cosmopedia-v2 (1M rows) fineweb-edu-dedup (1M rows) python-edu (0.5M rows) finemath-4plus (1M rows) tabular1M<n<10M2 likes452 downloads2y agoHugging Face08enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes340 downloads8mo agoHugging Face09enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes259 downloads8mo agoHugging Face10oieieio /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/oieieio/smollm-corpus.tabular100M<n<1B0 likes226 downloads1y agoHugging Face11david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes170 downloads4mo agoHugging Face12Mgmgrand420 /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/smollm-corpus.tabular100M<n<1B0 likes77 downloads8mo agoHugging Face13styal /very-smollm-corpus-0.5Mtabular100K<n<1M0 likes69 downloads2y agoHugging Face14ecreeth /1b-smollm-corpus SmolLM-Corpus — 1B Token Subset A curated 1-billion-token English pretraining corpus sampled from HuggingFaceTB/smollm-corpus, designed for training small language models (~20M parameters). Dataset Composition Source Ratio Tokens Documents FineWeb-Edu (dedup) 87% ~870M 849,577 Cosmopedia v2 13% ~130M 161,889 Total 100% ~1B 1,011,466 Rationale for the Split The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page: https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.texttext-generation1M<n<10M1 likes60 downloads4mo agoHugging Face15hhoenjet /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/hhoenjet/smollm-corpus.tabular100M<n<1B0 likes56 downloads7mo agoHugging Face16mesolitica /smollm-corpus-filter-malaysian-context HuggingFaceTB/smollm-corpus filter Malaysian context Originally from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus, source code at https://github.com/mesolitica/malaysian-dataset/tree/master/corpus/smollm-corpus we filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords using r5.16xlarge EC2 instance to filters 0 likes47 downloads2y agoHugging Face17Jakemu /smollm-corpus-shuffletext1M<n<10M0 likes46 downloads8mo agoHugging Face18Jakemu /smollm-corpus-and-FineWeb2-ko-synth-shuffle-20btext10M<n<100M0 likes46 downloads7mo agoHugging Face19OpenMOSS-Team /MHA2MLA-corpus-smollm_v10 likes31 downloads2y agoHugging Face20david-thrower /HelixLM-s512-instruct-smollm-stock-corpus-1-7m-token david-thrower/HelixLM-s512-instruct-smollm-stock-corpus-1-7m-token Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/HelixLM-s512-instruct-smollm-stock-corpus-1-7m-token.text1K<n<10K0 likes17 downloads4mo agoHugging Face21Obscure-Entropy /smollm-corpus-hutext100K<n<1M0 likes14 downloads1y agoHugging Face22eliebak /very-smollm-corpustext1M<n<10M3 likes13 downloads2y agoHugging Face23BEE-spoke-data /smollm-corpus-pythongated smollm-corpus - python A version of the python-edu subset with the text added tabulartext-generation10M<n<100M0 likes6 downloads9mo agoHugging Face24jsanzolac /smollm_corpustext1M<n<10M0 likes4 downloads5mo agoHugging Face25Jakemu /smollm-corpus-and-FineWeb2-ko-synth-shuffletext1M<n<10M0 likes3 downloads7mo agoHugging Face26RuiTenkawa /smollm-corpus-chunk0 likes2 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.