CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kenhktsui /cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia tabular10M<n<100M0 likes1.9k downloads2y agoHugging Face02BEE-spoke-data /cosmopedia-v2-mincols cosmopedia-v2: mincols cosmopedia-v2 with extra cols dropped to make the dataset smaller/easier to use texttext-generation10M<n<100M3 likes402 downloads9mo agoHugging Face03schuler /cosmopedia-v2-textbook-and-howto-8.3m Cosmopedia V2 Textbook and WikiHow Dataset 8.3M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-8.3m.texttext-generation1M<n<10M5 likes324 downloads2y agoHugging Face04enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes288 downloads8mo agoHugging Face05Lyric1010 /cosmopedia-v2-4B Dataset: cosmopedia-v2-4B This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-4B/stage_1/tmp. textn<1K0 likes225 downloads8mo agoHugging Face06schuler /cosmopedia-v2-textbook-and-howto-4.5m Cosmopedia V2 Textbook and WikiHow Dataset 4.5M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-4.5m.texttext-generation1M<n<10M0 likes222 downloads2y agoHugging Face07david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes171 downloads4mo agoHugging Face08kdcyberdude /cosmopedia_web_samples_v2_shards_entabular1M<n<10M0 likes135 downloads2y agoHugging Face09Hoglet-33 /CosmopediaV2-2BTThis dataset contains about 2 billion tokens from HuggingFaceTB/smollm-corpus from the Cosmopedia v2 subset. text1M<n<10M1 likes126 downloads2mo agoHugging Face10PatrickHaller /cosmopedia-v2-1BSubsample of smollm-corpus cosmopedia-v2 subset. Used following logic for sampling: for example in train_dataset: text = example['text'] if not text or not isinstance(text, str): continue num_tokens = count_tokens(text, tokenizer) remaining_tokens = target_tokens - total_tokens if remaining_tokens <= 0: break prob = min(1.0, remaining_tokens / (target_tokens * 0.1)) if random.random() < prob:… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cosmopedia-v2-1B.text1M<n<10M0 likes102 downloads2y agoHugging Face11Lyric1010 /cosmopedia-v2-noeod_8b Dataset: cosmopedia-v2-noeod_8b This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_8b/stage_1/tmp/. text0 likes87 downloads6mo agoHugging Face12Lyric1010 /cosmopedia-v2-noeod_4b Dataset: cosmopedia-v2-noeod_4b This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_4b/stage_1/tmp. text0 likes77 downloads8mo agoHugging Face13schuler /cosmopedia-v2-textbook-and-howto-2.3m Cosmopedia V2 Textbook and WikiHow Dataset 2.3M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-2.3m.texttext-generation1M<n<10M2 likes55 downloads2y agoHugging Face14semran1 /cosmopedia-v2-subsettext1M<n<10M0 likes50 downloads2y agoHugging Face15Dhiraj45 /cosmopedia-v2text1M<n<10M1 likes29 downloads5mo agoHugging Face16adorkin /cosmopedia-v2-translate-append-instructions-ettexttranslation1K<n<10K0 likes28 downloads1y agoHugging Face17audibeal74 /cosmopedia_v2tabular100K<n<1M0 likes11 downloads2mo agoHugging Face18Trelis /cosmopedia-v2-10percent-sampletext1M<n<10M0 likes8 downloads2y agoHugging Face19Lyric1010 /cosmopedia-v2-noeod_10b Dataset: cosmopedia-v2-noeod_10b This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_10b/stage_1. textn<1K0 likes4 downloads5mo agoHugging Face20Lyric1010 /cosmopedia-v2_10b_qwen3 Dataset: cosmopedia-v2_10b_qwen3 This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_qwen3/cosmopedia-v2_10b_qwen3/stage_1. text0 likes4 downloads5mo agoHugging Face21motionlabs /cosmopedia-v2-ranked-samplestext100K<n<1M0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.