CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajibawa-2023 /Children-Stories-CollectionChildren Stories Collection A great synthetic datasets consists of around 0.9 million stories especially meant for Young Children. You can directly use these datasets for training large models. Total 10 datasets are available for download. You can use any one or all the json files for training purpose. These datasets are in "prompt" and "text" format. Total token length is also available. Thank you for your love & support. texttext-generation100K<n<1M58 likes540 downloads3y agoHugging Face02ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes211 downloads9d agoHugging Face03ajibawa-2023 /Education-Young-ChildrenDetails coming soon!! text100K<n<1M16 likes104 downloads2y agoHugging Face04PinkPixel /Childrens-Story-Writing 🧒 Children's Story Writing Dataset ✨ This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences). 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_children.jsonl: Training data. writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.texttext-generation1M<n<10M2 likes67 downloads5mo agoHugging Face05nassimjp /pashto-reasoning-children-story-crafting-dataset Pashto Reasoning Children Story Crafting Dataset Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps. Dataset Overview & Methodology Language: Pashto (ps) Base Prompts: 100 unique core story prompts. Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.texttext-generationn<1K0 likes48 downloads5d agoHugging Face06nassimjp /stories_oh_children_pashtotext1K<n<10K0 likes23 downloads2mo agoHugging Face07nimuezorro /bilingual_children_speech Bilingual Children Speech Dataset Dataset Details This dataset was created with data taken from the Kaggle Dataset Corpus of bilingual children's speech. The original dataset includes much more data but for the purpose of this dataset only child utterances, l1, child_id, and age were extracted. The original dataset also includes much more free flowing dialogue and shorter utterances. Therefore, a script was used to extract the target child's English utterances and turn… See the full description on the dataset page: https://huggingface.co/datasets/nimuezorro/bilingual_children_speech.texttext-classification1K<n<10K0 likes20 downloads5mo agoHugging Face08thiemcun203 /Top_500_unsafe_for_children_prompttabularn<1K1 likes18 downloads2y agoHugging Face09asoria /children-stories-dataset children-stories-dataset Note: This is an AI-generated dataset, so its content may be inaccurate or false. Source of the data: The dataset was generated using Fastdata library and claude-3-haiku-20240307 with the following input: System Prompt You are a helpful assistant. Prompt Template Generate Children's Stories with title, content and the corresponding habit on the following topic <topic>{text}</topic> Sample Input {'idx': [0, 1], 'text':… See the full description on the dataset page: https://huggingface.co/datasets/asoria/children-stories-dataset.textn<1K1 likes11 downloads2y agoHugging Face10tasksource /children-tomtextn<1K0 likes6 downloads4y agoHugging Face11nikister /children-assistanttext10K<n<100K1 likes4 downloads1y agoHugging Face12Gabriel34190 /qcm-maths-childrenstextn<1K0 likes4 downloads1y agoHugging Face13arsalanaa /children_story_datasettext10K<n<100K0 likes1 downloads2y agoHugging Face14nassimjp /pashto-children-storiestext10K<n<100K0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.