datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Children-Stories-CollectionChildren Stories Collection
A great synthetic datasets consists of around 0.9 million stories especially meant for Young Children. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thank you for your love & support.
Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.Education-Young-ChildrenDetails coming soon!!
Childrens-Story-Writing
🧒 Children's Story Writing Dataset ✨
This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences).
📂 Dataset Structure
The data is provided in ChatML format, making it ideal for instruction tuning.
Files
writing_train_children.jsonl: Training data.
writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.stories_oh_children_pashtobilingual_children_speech
Bilingual Children Speech Dataset
Dataset Details
This dataset was created with data taken from the Kaggle Dataset Corpus of bilingual children's speech.
The original dataset includes much more data but for the purpose of this dataset only child utterances, l1, child_id, and age were extracted.
The original dataset also includes much more free flowing dialogue and shorter utterances.
Therefore, a script was used to extract the target child's English utterances and
turn… See the full description on the dataset page: https://huggingface.co/datasets/nimuezorro/bilingual_children_speech.Top_500_unsafe_for_children_promptchildren-stories-dataset
children-stories-dataset
Note: This is an AI-generated dataset, so its content may be inaccurate or false.
Source of the data:
The dataset was generated using Fastdata library and claude-3-haiku-20240307 with the following input:
System Prompt
You are a helpful assistant.
Prompt Template
Generate Children's Stories with title, content and the corresponding habit on the following topic <topic>{text}</topic>
Sample Input
{'idx': [0, 1], 'text':… See the full description on the dataset page: https://huggingface.co/datasets/asoria/children-stories-dataset.children-tomchildren-assistantqcm-maths-childrenschildren_story_datasetpashto-children-stories
