datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Children-Stories-CollectionChildren Stories Collection
A great synthetic datasets consists of around 0.9 million stories especially meant for Young Children. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thank you for your love & support.
Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.Childrens-Story-Writing
🧒 Children's Story Writing Dataset ✨
This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences).
📂 Dataset Structure
The data is provided in ChatML format, making it ideal for instruction tuning.
Files
writing_train_children.jsonl: Training data.
writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.Children-Stories-CollectionChildren Stories Collection
A great synthetic datasets consists of around 0.9 million stories especially meant for Young Children. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thank you for your love & support.
Children-Stories-Collection-Italian
Dataset Card for Dataset Name
Italian translation of https://huggingface.co/datasets/ajibawa-2023/Children-Stories-Collection
NOTE: I have decided to mantain alive the repo to honor the Colab CPU and Google GPU that work for it, but this dataset is rubbish. All stories are about tech interests (e.g., "Explain how to use PostSQL to kids, and how beatiful it is") or political questions total irrilevant to any kids.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/markod0925/Children-Stories-Collection-Italian.
