CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lyric1010 /cosmopedia-v2-4B Dataset: cosmopedia-v2-4B This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-4B/stage_1/tmp. textn<1K0 likes290 downloads8mo agoHugging Face02enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes230 downloads8mo agoHugging Face03agentlans /cosmopedia-classificationtexttext-classification100K<n<1M0 likes80 downloads29d agoHugging Face04aixsatoshi /cosmopedia-japanese-100kcosmopedia-japanese-20kのデータに、kunishou様から20k-100kをご提供いただけることになり100kまで拡大しました。   https://huggingface.co/datasets/kunishou/cosmopedia-100k-ja-preview   テキスト生成プロンプトの翻訳も含むデータは、上記レポジトリを確認してください。 text10K<n<100K7 likes46 downloads3y agoHugging Face05aixsatoshi /cosmopedia-japanese-20ktext10K<n<100K7 likes42 downloads3y agoHugging Face06aixsatoshi /Chat-with-cosmopediaReasoning、知識、会話の掛け合いなどの情報密度が高いマルチターンの会話データです。 利用データセット日本語化cosmopedia日本語化cosmopedia から作成した合成データセットです。 Example1 ユーザー: 数学をもっと身近に感じるためには、どのような取り組みが必要でしょうか? アシスタント: 数学を身近に感じるためには、適切な年齢層に合わせた教材やビデオ記録を利用することが効果的です。たとえば、MoMathなどの組織は、年齢に応じたコンテンツと戦略的なビデオ記録を利用することで、数学への参加を阻む障壁を取り除いています。これにより、STEM分野への幅広い参加が可能になり、かつてはエリート主義的なトピックを生み出すことで、将来の発見と革新の肥沃な土壌を作り出すことができます。 ユーザー: ビデオ記録がなぜ数学教育において重要なのでしょうか? アシスタント:… See the full description on the dataset page: https://huggingface.co/datasets/aixsatoshi/Chat-with-cosmopedia.text1K<n<10K3 likes42 downloads2y agoHugging Face07kunishou /cosmopedia-100k-ja-previewcosmopedia-100k のindex 20k ~ 100k を日本語に自動翻訳したデータになります(テキストが長すぎて翻訳エラーになったレコードは除外しています)。このデータセット自体は別作業者が取り組んでいる 0 ~ 20k の翻訳結果とマージ後に削除します。データセット自体は残しておきます。 text10K<n<100K4 likes37 downloads3y agoHugging Face08CJJones /Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 🖥️ Demo Interface: Discord Discord: https://discord.gg/Xe9tHFCS9h **Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.tabulartext-generation10K<n<100K2 likes15 downloads7mo agoHugging Face09Lyric1010 /cosmopedia-v2-noeod_10b Dataset: cosmopedia-v2-noeod_10b This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/cosmopedia-v2-noeod_10b/stage_1. textn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.