datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia
cosmopedia-v2-mincols
cosmopedia-v2: mincols
cosmopedia-v2 with extra cols dropped to make the dataset smaller/easier to use
cosmopedia-v2-textbook-and-howto-8.3m
Cosmopedia V2 Textbook and WikiHow Dataset 8.3M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-8.3m.smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.cosmopedia-v2-4B
Dataset: cosmopedia-v2-4B
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-4B/stage_1/tmp.
cosmopedia-v2-textbook-and-howto-4.5m
Cosmopedia V2 Textbook and WikiHow Dataset 4.5M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-4.5m.smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.cosmopedia_web_samples_v2_shards_enCosmopediaV2-2BTThis dataset contains about 2 billion tokens from HuggingFaceTB/smollm-corpus from the Cosmopedia v2 subset.
cosmopedia-v2-1BSubsample of smollm-corpus cosmopedia-v2 subset.
Used following logic for sampling:
for example in train_dataset:
text = example['text']
if not text or not isinstance(text, str):
continue
num_tokens = count_tokens(text, tokenizer)
remaining_tokens = target_tokens - total_tokens
if remaining_tokens <= 0:
break
prob = min(1.0, remaining_tokens / (target_tokens * 0.1))
if random.random() < prob:… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cosmopedia-v2-1B.cosmopedia-v2-noeod_8b
Dataset: cosmopedia-v2-noeod_8b
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_8b/stage_1/tmp/.
cosmopedia-v2-noeod_4b
Dataset: cosmopedia-v2-noeod_4b
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_4b/stage_1/tmp.
cosmopedia-v2-textbook-and-howto-2.3m
Cosmopedia V2 Textbook and WikiHow Dataset 2.3M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-2.3m.cosmopedia-v2-subsetcosmopedia-v2cosmopedia-v2-translate-append-instructions-etcosmopedia_v2cosmopedia-v2-10percent-samplecosmopedia-v2-noeod_10b
Dataset: cosmopedia-v2-noeod_10b
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_10b/stage_1.
cosmopedia-v2_10b_qwen3
Dataset: cosmopedia-v2_10b_qwen3
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_qwen3/cosmopedia-v2_10b_qwen3/stage_1.
cosmopedia-v2-ranked-samples
