CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SimpleStories /SimpleStories 📘📕 SimpleStories 📙📗 SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository. If you'd like to commission other languages or story formats, feel free to send mail. When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.tabulartext-generation1M<n<10M39 likes2.9k downloads9mo agoHugging Face02SimpleStories /SimpleStories-JA 📘📕 SimpleStories 📙📗 このデータセットは、gpt-4o-miniによって生成された短編小説で出来ているデータセットです。生成方法や、自分で物語を生成する方法については、こちらのリポジトリをご覧ください。 他の言語や物語形式の制作を希望される場合は、メールにてお問い合わせください。 SimpleStoriesは、EldenとLiによるTinyStoriesの改良版です。 特徴 物語の注釈情報(theme、topic、styleなど) 多様性の高さ 2024年のモデルによって生成 NLPのデータが用意しているためフィルタリングしやすい 以下の言語版が利用可能: 英語 日本語 他にも追加予定 This dataset is a collection of short stories generated by gpt-4o-mini (+ other models, soon). To see how this dataset was generated, or to generate some stories… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories-JA.tabulartext-generation1M<n<10M1 likes137 downloads2y agoHugging Face03duoduoyeah /SimpleStories 📘📕 SimpleStories 📙📗 SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository. If you'd like to commission other languages or story formats, feel free to send mail. When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/SimpleStories.tabulartext-generation1M<n<10M0 likes66 downloads9mo agoHugging Face04keenanpepper /simplestories-ja-en-endings Dataset Card for SimpleStories-JA-EN-Endings I translated the endings of all the stories in SimpleStories-JA into English using Claude 3.5 Haiku. There was over a 99% success rate so the number of stories in here is just slightly less than the number in SimpleStories-JA. Dataset Details Dataset Sources SimpleStories/SimpleStories-JA texttext-generation1M<n<10M0 likes38 downloads1y agoHugging Face05trixyL /simplestories-8k-megatron📦 Megatron-LM/MegaDLMs Preprocessed Dataset This dataset hosts the Megatron-LM/MegaDLMs preprocessed SimpleStories dataset using an 8k vocab BPE Tokenizer. ✅ What this contains Preprocessed Megatron dataset files (e.g., .bin / .idx) ready for Megatron-LM/MegaDLMs training BPE Tokenizer config files used to create the dataset 🔗 References Fork with extra preprocessing utils for SimpleStories: https://github.com/triloy8/MegaDLMs Original MegaDLMs repo:… See the full description on the dataset page: https://huggingface.co/datasets/trixyL/simplestories-8k-megatron.text-generation1M<n<10M1 likes31 downloads8mo agoHugging Face06SmallScale /Simple-Stories-Hindi Hindi Simple Stories Overview Hindi Simple Stories is a Hindi translation of the Simple Stories dataset. It is designed as a clean corpus of simple, coherent stories for training and evaluating Hindi language models. The dataset preserves the structure and meaning of the original stories while making them available in Hindi for language modeling research. Motivation Despite rapid progress in large language models, there are relatively few… See the full description on the dataset page: https://huggingface.co/datasets/SmallScale/Simple-Stories-Hindi.text-generation2 likes29 downloads2mo agoHugging Face07SmallScale /Simple-Stories-Eng-Hin Simple-Stories-Eng-Hin A parallel English–Hindi dataset of short, simple stories. Each row contains the same story written in English and its Hindi counterpart, making it useful for training or evaluating small language models, translation systems, and Hindi text generation models on simple narrative text. Dataset Details Rows: ~1.72M Format: JSON (auto-converted to Parquet) Split: train (single split) Size: ~9.38 GB License: MIT Columns… See the full description on the dataset page: https://huggingface.co/datasets/SmallScale/Simple-Stories-Eng-Hin.text-generation1M<n<10M2 likes16 downloads2mo agoHugging Face08trixyL /simplestories-4k-megatron📦 Megatron-LM/MegaDLMs Preprocessed Dataset This dataset hosts the Megatron-LM/MegaDLMs preprocessed SimpleStories dataset using an 4k vocab BPE Tokenizer. ✅ What this contains Preprocessed Megatron dataset files (e.g., .bin / .idx) ready for Megatron-LM/MegaDLMs training BPE Tokenizer config files used to create the dataset 🔗 References Fork with extra preprocessing utils for SimpleStories: https://github.com/triloy8/MegaDLMs Original MegaDLMs repo:… See the full description on the dataset page: https://huggingface.co/datasets/trixyL/simplestories-4k-megatron.text-generation1M<n<10M1 likes13 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.