datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStoriesV2Data is from https://huggingface.co/datasets/roneneldan/TinyStories/:
TinyStoriesV2-GPT4-train.txt - Is a new version of the dataset that is based on generations by GPT-4 only (the original dataset also has generations by GPT-3.5 which are of lesser quality). It contains all the examples in TinyStories.txt which were GPT-4 generated as a subset (but is significantly larger).
This dataset was used to train https://github.com/noanabeshima/tiny_model/.
The data was preprocessed with:
from… See the full description on the dataset page: https://huggingface.co/datasets/noanabeshima/TinyStoriesV2.somali-tinystoriestinystories-icr-data-v2tinystories_dataset_arabicTinyStories-Multilingual
Novelist: TinyStories Multilingual Edition
Dataset Summary
The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes.
The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.TinyStories-SFT
This is Finetuning dataset for tinystories
This is for finetuning model trained on dataset https://huggingface.co/datasets/roneneldan/TinyStories
json on this dataset
I'll update readme later
Tiny-Ko-Stories
Tiny-Ko-Stories
English version is available below.
Tiny-Ko-Stories는 TinyStories에서 영감을 받은 한국어 이야기 데이터셋입니다.
TinyStories는 제한된 고품질 데이터셋을 사용하면, 소형 모델이라도 추론 능력과 창의력을 발휘할 수 있음을 보였습니다.
우리가 확인하려는 것은 단순합니다.
이 현상이 한국어에서도 재현될까?
이를 확인하려면 번역 데이터셋만으로는 부족했습니다. 한국어다운 이름, 문장 리듬, 의성어와 의태어, 색채어, 작은 사건 구조를 포함하려면 처음부터 한국어로 만든 이야기가 필요했습니다. 그래서 Tiny-Ko-Stories는 영어 TinyStories를 번역하는 대신, 한국어로 짧은 이야기를 새로 생성하고 여러 단계의 검수를 거쳐 구성했습니다.
데이터셋 요약
항목
값
레코드 수
2,003,542
형식
JSONL
공개… See the full description on the dataset page: https://huggingface.co/datasets/psymon/Tiny-Ko-Stories.TinyStories_igbo
TinyStories English-Igbo Parallel Corpus
Description
The translations were derived from the original English TinyStories dataset. Each story has been carefully translated to retain the simplicity and educational value intended for the target age group. The datasets are organized into two main files:
Composition
Igbo Translations: Contains stories translated into Igbo, with the original English texts for reference.
Each file is structured to include the… See the full description on the dataset page: https://huggingface.co/datasets/ccibeekeoc42/TinyStories_igbo.TinyStories-GPT4-V2-50K-SUBSETTinyStories_yoruba
TinyStories English-Igbo Parallel Corpus
Description
TBD
Composition
TBD
Usage
TBD
Acknowledgments
TBD
License
The translated datasets are released under Apache2.0, consistent with the original TinyStories dataset's licensing terms. Please refer to Microsoft's official release for further details on the licensing of the TinyStories dataset.
About the Authors
Christopher Ibe and Okezie Okoye continue to lead Hypa AI towards new… See the full description on the dataset page: https://huggingface.co/datasets/ccibeekeoc42/TinyStories_yoruba.TinyStories-GPT4-V2marathi-generated_4o-mini_2MTinyStories-translate-4k원래는 데이터셋 전체를 번역해서 올리는 것이 목표였지만 문제가 너무 많이 터진 탓에 이거라도 올립니다
tiny-stories-deTinyStories-korean-eduTinyStories와 TinyStories-Korean, fineweb-edu-classifier를 이용하여 만든 데이터입니다.
TinyStories를 fineweb-edu-classifier로 평가를 한 뒤 int_score가 3 이상이면 이에 해당되는 TinyStories-Korean을 데이터에 추가하는 방식으로 제작하였습니다.
TinyStories-Korean에서 데이터 300개가 누락되어 있어서 데이터 누락이 되지 않음을 확인한 1832000번째 데이터까지만 평가하였습니다.
score 비율
3: 99.4%(501,651개)
4: 0.6%(2,931개)
beng-generated_4o-mini_2Mmarathi-tinystories-10khindi-generated_4o-mini_2Mmorpeheme_processed_tinystoriesThis is a modified version of a Turkish TinyStories dataset: https://huggingface.co/datasets/umarigan/tinystories_tr
The modification is that suffixes are replaced with PUA (Private Use Area) characters. Suffixes are preceded by a label that specifies whether the word is a noun, verb, or named entity.
Here is a full list of all suffixes and word labels:
[
"Root", "Noun", "Adj", "Verb", "Pron", "Adv", "Conj", "Punc", "Ques",
"Postp", "Det", "Num", "Dup", "Interj", "A1sg", "A2sg"… See the full description on the dataset page: https://huggingface.co/datasets/esat-krky/morpeheme_processed_tinystories.Tiny-Stories-1500hindi-generated_4o-mini_2MTinyStories-shuffled
Shuffled Dataset
This dataset was created by shuffling hynky/elon_tweets using merge-shuffle.
All original columns are preserved — rows are output as JSONL.
Parameters
Parameter
Value
Source
hynky/elon_tweets
Format
JSONL (all columns)
Seed
42
Buckets
256
Output size
0.00 GB
Shards
1
Shuffle time
1.0s
Upload time
2.8s
TinyStoriesV2-GPT4marathi-generated_4o-mini_2Mbangla-generated_4o-mini_2MTinyStoriesInFrench
