datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.SimpleStories-JA
📘📕 SimpleStories 📙📗
このデータセットは、gpt-4o-miniによって生成された短編小説で出来ているデータセットです。生成方法や、自分で物語を生成する方法については、こちらのリポジトリをご覧ください。
他の言語や物語形式の制作を希望される場合は、メールにてお問い合わせください。
SimpleStoriesは、EldenとLiによるTinyStoriesの改良版です。
特徴
物語の注釈情報(theme、topic、styleなど)
多様性の高さ
2024年のモデルによって生成
NLPのデータが用意しているためフィルタリングしやすい
以下の言語版が利用可能:
英語
日本語
他にも追加予定
This dataset is a collection of short stories generated by gpt-4o-mini (+ other models, soon). To see how this dataset was generated, or to generate some stories… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories-JA.SimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/SimpleStories.simplestories-ja-en-endings
Dataset Card for SimpleStories-JA-EN-Endings
I translated the endings of all the stories in SimpleStories-JA into English using Claude 3.5 Haiku.
There was over a 99% success rate so the number of stories in here is just slightly less than the number in SimpleStories-JA.
Dataset Details
Dataset Sources
SimpleStories/SimpleStories-JA
simplestoriesplussimplestories-persona-clusters-augmentsimplestories-personas
SimpleStories Personas
5 stylistic-persona variants of the SimpleStories corpus, each ~2150 short
stories generated by gpt-4o-mini with style instructions specific to the
persona. Used to study mixture-of-personas fine-tuning behavior.
Personas
noir_detective — first-person hardboiled detective narrator; short, clipped sentences; dark atmosphere.
fairy_tale — classic "Once upon a time" opener; third-person omniscient; ends with an explicit moral.
scientific_explainer —… See the full description on the dataset page: https://huggingface.co/datasets/desh2806/simplestories-personas.simplestoriesplus-samplesimple_user_storiesSimpleStories-persona-augmentedsimple_storiesedition_1575_SimpleStories-SimpleStories-readymade
edition_1575_SimpleStories-SimpleStories-readymade
A Readymade by TheFactoryX
Original Dataset
SimpleStories/SimpleStories
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1575_SimpleStories-SimpleStories-readymade.simple_stories_newSimpleStories-clusteredsimplestories-personas-10k
