datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Children_Image_Captioning_Dataset_Split0
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.Agilex_Cobot_Magic_fold_jeans_shorts_children_s
Agilex_Cobot_Magic_fold_jeans_shorts_children's
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 50
Total Frames: 60562
FPS: 30
Dataset Size: 912.51 MB
Robot Name: Agilex_Cobot_Magic
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_fold_jeans_shorts_children_s.samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.Children-Stories-CollectionChildren Stories Collection
A great synthetic datasets consists of around 0.9 million stories especially meant for Young Children. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thank you for your love & support.
children-story-datasetG1_WBT_Dex1_Building-Children-Table
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (2)
Finger states for both hands.
Dex1 Hand (range: 5.5 – 0.0, open → close)
Per hand:
Open/close… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/G1_WBT_Dex1_Building-Children-Table.Chinese_Children_Image_Captioning_Dataset_Split1
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.Children-zhAudio-Children-Stories-CollectionAudio Chidren Stories Collection
This dataset has 600 audio files in .mp3 format. This has been created using my existing dataset Children-Stories-Collection.
I have used only first 600 stories for creating this audio dataset.
You can use this for training and research purpose.
Thank you for your love & support.
health-conditions-among-children-under-age-18-by-s
Health conditions among children under age 18, by selected characteristics: United States
Description
NOTE: On October 19, 2021, estimates for 2016–2018 by health insurance status were revised to correct errors. Changes are highlighted and tagged at https://www.cdc.gov/nchs/data/hus/2019/012-508.pdf
Data on health conditions among children under age 18, by selected population characteristics. Please refer to the PDF or Excel version of this table in the HUS 2019 Data… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/health-conditions-among-children-under-age-18-by-s.africa-who-distribution-of-causes-of-death-among-children-aged-5-years
Africa — WHO GHO: Distribution of causes of death among children aged < 5 years (%) | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-distribution-of-causes-of-death-among-children-aged-5-years.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.Audio-Children-Stories-Collection-LargeAudio Chidren Stories Collection Large
This dataset has 5600++ audio files in .mp3 format. This has been created using my existing dataset Children-Stories-Collection.
I have used first 5600++ stories from Children-Stories-1-Final.json file for creating this audio dataset.
You can use this for training and research purpose.
Thank you for your love & support.
Synthetic-Childrens-StoriesThis dataset has moved to ContextReq/Synthetic-Dataset-Childrens-Stories → https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories
Children_Intent_Classification
MAMA Communicative Intent Dataset (INCA-A Annotated)
Overview
The MAMA Communicative Intent Dataset is a linguistically annotated corpus of child utterances designed to support research in child-centred Natural Language Processing (NLP) and communicative intent recognition in early language development.
The dataset contains 10,800 child utterances annotated using the INCA Communicative Coding System (Ninio et al., 1994), a developmental framework that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/Wajinimi/Children_Intent_Classification.G1_WBT_Dex1_Building-Children-Table-SONICEducation-Young-ChildrenDetails coming soon!!
clothes_for_men_women_children
Image Description Dataset
Dataset Description
This dataset contains 3082 images with their corresponding descriptions in both long and short formats.
The descriptions were generated using the BLIP-large model.
Dataset Statistics
Total images: 3082
Average words in long description: 17.5
Average words in short description: 8.8
Languages
English (en)
Dataset Structure
Each record in the dataset contains:
file_name: Relative path to the… See the full description on the dataset page: https://huggingface.co/datasets/AntZet/clothes_for_men_women_children.ZJU-Children-Emotion
ZJU Children Emotion Dataset
Dataset Description
ZJU Children Emotion Dataset (ZCED) is a multi-modal multi-group children emotion dataset aimed for the research of disease classification and emotion recognition in children with neurodevelopmental disorders.
The ZCED dataset contains both behavioral and physiological recordings based on a video-based emotional stimulation paradigm for four groups of children including 19 TD, 15 ASD, 20 ADHD, and 18 ASD+ADHD.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/Jiaheng-Wang/ZJU-Children-Emotion.men_women_children_wearing_clothes
Image Description Dataset
Dataset Description
This dataset contains 6979 images with their corresponding descriptions in both long and short formats.
The descriptions were generated using the BLIP-large model.
Dataset Statistics
Total images: 6979
Average words in long description: 17.3
Average words in short description: 9.4
Languages
English (en)
Dataset Structure
Each record in the dataset contains:
file_name: Relative path to the… See the full description on the dataset page: https://huggingface.co/datasets/AntZet/men_women_children_wearing_clothes.the40thai-children-stories-tha-classification
The40ThaiChildrenStories_tha_Classification
Deduplicated copy of kornwtp/the40thai-children-stories-tha-classification.
Splits
split
rows
train
1,950
G1_WBT_Dex1_Building-Children-Table-SONIC-10episodeshar_children_2024-harth
Tørring 2024 — thigh + back accelerometry, children (typically developing + cerebral palsy), activity recognition
Dual-sensor accelerometry (Axivity AX3, 50 Hz, ±8 g) from the thigh and lower back
in children with and without cerebral palsy, with activity ground truth for 13 activity
types including walking, running, jumping, cycling, and postural activities. Recordings
were collected across lab, gymnasium, and outdoor settings. Harmonized from the Dataverse
release into… See the full description on the dataset page: https://huggingface.co/datasets/josefheidler/har_children_2024-harth.children-stories-dataset
Children's Stories Dataset
Dataset Description
This dataset contains a collection of children's stories designed to teach positive values, problem-solving skills, and emotional intelligence. The stories feature diverse characters and settings, making them suitable for children aged 3-8.
Dataset Structure
Each story record contains:
id: Unique story identifier
title: Story title
text: Full story text
type: Story type (e.g., "daily_adventure"… See the full description on the dataset page: https://huggingface.co/datasets/garethpaul/children-stories-dataset.indonesian-children-news
Indonesian Children News Dataset
This dataset contains articles from Indonesian children's news sources.
Dataset Description
Dataset Statistics
Number of articles: 18,507
Total number of tokens: 9,531,263
Number of unique tokens: 85,341
Average tokens per article: 515.01
The dataset contains these columns:
text: The full text content of the article
title: The article title
url: Source URL
published_date: Article publication date
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/haznitrama/indonesian-children-news.children-faces-synthetic-deepfake-detection
Dataset of synthetic kids faces
Children are missing in most deepfake prevention and age detection datasets. This underrepresentation leads to biased models, weaker age-restricted access controls, and increased vulnerability to deepfake attacks. The Synthetic Children Faces Dataset fills that gap: it expands demographic coverage while avoiding any real-person data
Dataset Features
Dataset Size: Thousands of high-quality images across 5–9, 9–12, 12–16 age groups… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/children-faces-synthetic-deepfake-detection.Childrens-Story-Writing
🧒 Children's Story Writing Dataset ✨
This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences).
📂 Dataset Structure
The data is provided in ChatML format, making it ideal for instruction tuning.
Files
writing_train_children.jsonl: Training data.
writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.asia-who-stunting-prevalence-among-children-under-5-years-of-age-untingprev
Stunting prevalence among children under 5 years of age (% height-for-age <-2 SD), model-based estimates | Asia (WHO GHO)
🌏 3,375 observations · 45 Asia countries · 2000–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 3,375 observations of Stunting prevalence among children under 5 years of age (% height-for-age <-2 SD), model-based estimates data across 45 Asia countries, spanning 2000–2024, covering 1 distinct indicators.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-stunting-prevalence-among-children-under-5-years-of-age-untingprev.asia-who-overweight-numbers-among-children-under-5-years-of-age
Overweight numbers among children under 5 years of age (thousands), model-based estimates | Asia (WHO GHO)
🌏 3,300 observations · 44 Asia countries · 2000–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 3,300 observations of Overweight numbers among children under 5 years of age (thousands), model-based estimates data across 44 Asia countries, spanning 2000–2024, covering 1 distinct indicators.
About the source
Source:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-overweight-numbers-among-children-under-5-years-of-age.indonesian-children-books
Indonesian Children Books Dataset
This dataset contains text extracted from Indonesian children's books. This dataset is still contains raw text directly extracted from books, therefore still considered as dirty and need to be preprocessed further.
Dataset Description
Dataset Statistics
Number of books: 2,740
Total pages: 165,245
Total number of tokens: 25,759,439
Number of unique tokens: 698,094
Average tokens per page: 155.89
Extraction Methods… See the full description on the dataset page: https://huggingface.co/datasets/haznitrama/indonesian-children-books.
