datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStoriesV2Data is from https://huggingface.co/datasets/roneneldan/TinyStories/:
TinyStoriesV2-GPT4-train.txt - Is a new version of the dataset that is based on generations by GPT-4 only (the original dataset also has generations by GPT-3.5 which are of lesser quality). It contains all the examples in TinyStories.txt which were GPT-4 generated as a subset (but is significantly larger).
This dataset was used to train https://github.com/noanabeshima/tiny_model/.
The data was preprocessed with:
from… See the full description on the dataset page: https://huggingface.co/datasets/noanabeshima/TinyStoriesV2.somali-tinystoriestinystories-icr-data-v2tinystories_dataset_arabicTinyStories-Multilingual
Novelist: TinyStories Multilingual Edition
Dataset Summary
The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes.
The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.TinyStories-SFT
This is Finetuning dataset for tinystories
This is for finetuning model trained on dataset https://huggingface.co/datasets/roneneldan/TinyStories
json on this dataset
I'll update readme later
TinyStories_igbo
TinyStories English-Igbo Parallel Corpus
Description
The translations were derived from the original English TinyStories dataset. Each story has been carefully translated to retain the simplicity and educational value intended for the target age group. The datasets are organized into two main files:
Composition
Igbo Translations: Contains stories translated into Igbo, with the original English texts for reference.
Each file is structured to include the… See the full description on the dataset page: https://huggingface.co/datasets/ccibeekeoc42/TinyStories_igbo.TinyStories-GPT4-V2-50K-SUBSETTinyStories_yoruba
TinyStories English-Igbo Parallel Corpus
Description
TBD
Composition
TBD
Usage
TBD
Acknowledgments
TBD
License
The translated datasets are released under Apache2.0, consistent with the original TinyStories dataset's licensing terms. Please refer to Microsoft's official release for further details on the licensing of the TinyStories dataset.
About the Authors
Christopher Ibe and Okezie Okoye continue to lead Hypa AI towards new… See the full description on the dataset page: https://huggingface.co/datasets/ccibeekeoc42/TinyStories_yoruba.TinyStories-GPT4-V2marathi-generated_4o-mini_2MTinyStories-translate-4k원래는 데이터셋 전체를 번역해서 올리는 것이 목표였지만 문제가 너무 많이 터진 탓에 이거라도 올립니다
tiny-stories-deTinyStories-korean-eduTinyStories와 TinyStories-Korean, fineweb-edu-classifier를 이용하여 만든 데이터입니다.
TinyStories를 fineweb-edu-classifier로 평가를 한 뒤 int_score가 3 이상이면 이에 해당되는 TinyStories-Korean을 데이터에 추가하는 방식으로 제작하였습니다.
TinyStories-Korean에서 데이터 300개가 누락되어 있어서 데이터 누락이 되지 않음을 확인한 1832000번째 데이터까지만 평가하였습니다.
score 비율
3: 99.4%(501,651개)
4: 0.6%(2,931개)
beng-generated_4o-mini_2Mmarathi-tinystories-10khindi-generated_4o-mini_2Mmorpeheme_processed_tinystoriesThis is a modified version of a Turkish TinyStories dataset: https://huggingface.co/datasets/umarigan/tinystories_tr
The modification is that suffixes are replaced with PUA (Private Use Area) characters. Suffixes are preceded by a label that specifies whether the word is a noun, verb, or named entity.
Here is a full list of all suffixes and word labels:
[
"Root", "Noun", "Adj", "Verb", "Pron", "Adv", "Conj", "Punc", "Ques",
"Postp", "Det", "Num", "Dup", "Interj", "A1sg", "A2sg"… See the full description on the dataset page: https://huggingface.co/datasets/esat-krky/morpeheme_processed_tinystories.Tiny-Stories-1500hindi-generated_4o-mini_2MTinyStories-shuffled
Shuffled Dataset
This dataset was created by shuffling hynky/elon_tweets using merge-shuffle.
All original columns are preserved — rows are output as JSONL.
Parameters
Parameter
Value
Source
hynky/elon_tweets
Format
JSONL (all columns)
Seed
42
Buckets
256
Output size
0.00 GB
Shards
1
Shuffle time
1.0s
Upload time
2.8s
TinyStoriesV2-GPT4marathi-generated_4o-mini_2Mbangla-generated_4o-mini_2MTinyStoriesInFrench
