cosmopedia-v2
cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia
smollm-cosmopedia-v2cosmopedia-v2-mincols
cosmopedia-v2: mincols
cosmopedia-v2 with extra cols dropped to make the dataset smaller/easier to use
cosmopedia-v2-textbook-and-howto-8.3m
Cosmopedia V2 Textbook and WikiHow Dataset 8.3M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-8.3m.smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.cosmopedia-v2-4B
Dataset: cosmopedia-v2-4B
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-4B/stage_1/tmp.
