datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Children_Image_Captioning_Dataset_Split0
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.Children-Stories-CollectionChildren Stories Collection
A great synthetic datasets consists of around 0.9 million stories especially meant for Young Children. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thank you for your love & support.
Chinese_Children_Image_Captioning_Dataset_Split1
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.Children-zhhealth-conditions-among-children-under-age-18-by-s
Health conditions among children under age 18, by selected characteristics: United States
Description
NOTE: On October 19, 2021, estimates for 2016–2018 by health insurance status were revised to correct errors. Changes are highlighted and tagged at https://www.cdc.gov/nchs/data/hus/2019/012-508.pdf
Data on health conditions among children under age 18, by selected population characteristics. Please refer to the PDF or Excel version of this table in the HUS 2019 Data… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/health-conditions-among-children-under-age-18-by-s.africa-who-distribution-of-causes-of-death-among-children-aged-5-years
Africa — WHO GHO: Distribution of causes of death among children aged < 5 years (%) | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-distribution-of-causes-of-death-among-children-aged-5-years.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.Children_Intent_Classification
MAMA Communicative Intent Dataset (INCA-A Annotated)
Overview
The MAMA Communicative Intent Dataset is a linguistically annotated corpus of child utterances designed to support research in child-centred Natural Language Processing (NLP) and communicative intent recognition in early language development.
The dataset contains 10,800 child utterances annotated using the INCA Communicative Coding System (Ninio et al., 1994), a developmental framework that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/Wajinimi/Children_Intent_Classification.Education-Young-ChildrenDetails coming soon!!
clothes_for_men_women_children
Image Description Dataset
Dataset Description
This dataset contains 3082 images with their corresponding descriptions in both long and short formats.
The descriptions were generated using the BLIP-large model.
Dataset Statistics
Total images: 3082
Average words in long description: 17.5
Average words in short description: 8.8
Languages
English (en)
Dataset Structure
Each record in the dataset contains:
file_name: Relative path to the… See the full description on the dataset page: https://huggingface.co/datasets/AntZet/clothes_for_men_women_children.men_women_children_wearing_clothes
Image Description Dataset
Dataset Description
This dataset contains 6979 images with their corresponding descriptions in both long and short formats.
The descriptions were generated using the BLIP-large model.
Dataset Statistics
Total images: 6979
Average words in long description: 17.3
Average words in short description: 9.4
Languages
English (en)
Dataset Structure
Each record in the dataset contains:
file_name: Relative path to the… See the full description on the dataset page: https://huggingface.co/datasets/AntZet/men_women_children_wearing_clothes.the40thai-children-stories-tha-classification
The40ThaiChildrenStories_tha_Classification
Deduplicated copy of kornwtp/the40thai-children-stories-tha-classification.
Splits
split
rows
train
1,950
har_children_2024-harth
Tørring 2024 — thigh + back accelerometry, children (typically developing + cerebral palsy), activity recognition
Dual-sensor accelerometry (Axivity AX3, 50 Hz, ±8 g) from the thigh and lower back
in children with and without cerebral palsy, with activity ground truth for 13 activity
types including walking, running, jumping, cycling, and postural activities. Recordings
were collected across lab, gymnasium, and outdoor settings. Harmonized from the Dataverse
release into… See the full description on the dataset page: https://huggingface.co/datasets/josefheidler/har_children_2024-harth.indonesian-children-news
Indonesian Children News Dataset
This dataset contains articles from Indonesian children's news sources.
Dataset Description
Dataset Statistics
Number of articles: 18,507
Total number of tokens: 9,531,263
Number of unique tokens: 85,341
Average tokens per article: 515.01
The dataset contains these columns:
text: The full text content of the article
title: The article title
url: Source URL
published_date: Article publication date
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/haznitrama/indonesian-children-news.Childrens-Story-Writing
🧒 Children's Story Writing Dataset ✨
This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences).
📂 Dataset Structure
The data is provided in ChatML format, making it ideal for instruction tuning.
Files
writing_train_children.jsonl: Training data.
writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.asia-who-stunting-prevalence-among-children-under-5-years-of-age-untingprev
Stunting prevalence among children under 5 years of age (% height-for-age <-2 SD), model-based estimates | Asia (WHO GHO)
🌏 3,375 observations · 45 Asia countries · 2000–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 3,375 observations of Stunting prevalence among children under 5 years of age (% height-for-age <-2 SD), model-based estimates data across 45 Asia countries, spanning 2000–2024, covering 1 distinct indicators.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-stunting-prevalence-among-children-under-5-years-of-age-untingprev.asia-who-overweight-numbers-among-children-under-5-years-of-age
Overweight numbers among children under 5 years of age (thousands), model-based estimates | Asia (WHO GHO)
🌏 3,300 observations · 44 Asia countries · 2000–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 3,300 observations of Overweight numbers among children under 5 years of age (thousands), model-based estimates data across 44 Asia countries, spanning 2000–2024, covering 1 distinct indicators.
About the source
Source:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-overweight-numbers-among-children-under-5-years-of-age.indonesian-children-books
Indonesian Children Books Dataset
This dataset contains text extracted from Indonesian children's books. This dataset is still contains raw text directly extracted from books, therefore still considered as dirty and need to be preprocessed further.
Dataset Description
Dataset Statistics
Number of books: 2,740
Total pages: 165,245
Total number of tokens: 25,759,439
Number of unique tokens: 698,094
Average tokens per page: 155.89
Extraction Methods… See the full description on the dataset page: https://huggingface.co/datasets/haznitrama/indonesian-children-books.asia-who-underweight-prevalence-among-children-under-5-years-of-age
Underweight prevalence among children under 5 years of age (% weight-for-age <-2 SD), survey-based estimates | Asia (WHO GHO)
🌏 20,493 observations · 45 Asia countries · 1986–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 20,493 observations of Underweight prevalence among children under 5 years of age (% weight-for-age <-2 SD), survey-based estimates data across 45 Asia countries, spanning 1986–2024, covering 1 distinct indicators.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-underweight-prevalence-among-children-under-5-years-of-age.illustrations_for_children
数据集 README
数据集概述
欢迎使用我们的数据集,该数据集主要包含网络收集的儿童插画(儿插)。这些插画旨在为教育和研究目的提供丰富的视觉素材。我们鼓励用户在遵守本README中规定的条款和条件的前提下,充分利用这些资源进行学习和研究。
许可协议
本数据集遵循Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0)许可协议。这意味着您可以:
自由分享:复制和分发数据集中的材料。
自由改编:基于本数据集的材料进行修改和再创作。
但请注意以下限制:
非商业性:您不得将本数据集用于商业目的。
相同方式共享:如果您对数据集进行了修改或衍生,您必须以相同的许可协议分发您的作品。
署名:您必须给出适当的署名,提供许可协议链接,并说明是否进行了更改。您可以以任何合理的方式进行署名,但不得以任何方式暗示许可人认可您或您的使用。
使用限制… See the full description on the dataset page: https://huggingface.co/datasets/shiertier/illustrations_for_children.asia-who-number-of-deaths-among-children-ages-5-to-9-years
Number of deaths among children ages 5 to 9 years | Asia (WHO GHO)
🌏 4,896 observations · 48 Asia countries · 1990–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,896 observations of Number of deaths among children ages 5 to 9 years data across 48 Asia countries, spanning 1990–2023, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-number-of-deaths-among-children-ages-5-to-9-years.asia-who-overweight-prevalence-among-children-under-5-years-of-age-weightprev
Overweight prevalence among children under 5 years of age (% weight-for-height >+2 SD), model-based estimates | Asia (WHO GHO)
🌏 3,300 observations · 44 Asia countries · 2000–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 3,300 observations of Overweight prevalence among children under 5 years of age (% weight-for-height >+2 SD), model-based estimates data across 44 Asia countries, spanning 2000–2024, covering 1 distinct indicators.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-overweight-prevalence-among-children-under-5-years-of-age-weightprev.asia-who-stunting-numbers-among-children-under-5-years-of-age
Stunting numbers among children under 5 years of age (millions), model-based estimates | Asia (WHO GHO)
🌏 3,375 observations · 45 Asia countries · 2000–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 3,375 observations of Stunting numbers among children under 5 years of age (millions), model-based estimates data across 45 Asia countries, spanning 2000–2024, covering 1 distinct indicators.
About the source
Source: WHO… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-stunting-numbers-among-children-under-5-years-of-age.europe-who-stunting-prevalence-among-children-under-5-years-of-age-nanthazne2
Stunting prevalence among children under 5 years of age (% height-for-age <-2 SD), survey-based estimates | Europe (WHO GHO)
🇪🇺 4,065 observations · 24 Europe countries · 1991–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 4,065 observations of Stunting prevalence among children under 5 years of age (% height-for-age <-2 SD), survey-based estimates data across 24 Europe countries, spanning 1991–2023, covering 1 distinct indicators.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-who-stunting-prevalence-among-children-under-5-years-of-age-nanthazne2.asia-who-mortality-rate-among-children-ages-5-to-9-years
Mortality rate among children ages 5 to 9 years (per 1000 children aged 5) | Asia (WHO GHO)
🌏 4,896 observations · 48 Asia countries · 1990–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,896 observations of Mortality rate among children ages 5 to 9 years (per 1000 children aged 5) data across 48 Asia countries, spanning 1990–2023, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-mortality-rate-among-children-ages-5-to-9-years.samromur_children_test
Dataset Card for samromur_children
Dataset Summary
The Samrómur Children Corpus consists of audio recordings and metadata files containing prompts read by the participants. It contains more than 137000 validated speech-recordings uttered by Icelandic children.
The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab (LVL) at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology. The recording process has… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/samromur_children_test.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.asia-who-number-of-deaths-among-children-ages-10-to-14-years
Number of deaths among children ages 10 to 14 years | Asia (WHO GHO)
🌏 4,608 observations · 48 Asia countries · 1990–2021 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,608 observations of Number of deaths among children ages 10 to 14 years data across 48 Asia countries, spanning 1990–2021, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-number-of-deaths-among-children-ages-10-to-14-years.asia-who-estimated-number-of-children-needing-antiretroviral-therapy
Estimated number of children needing antiretroviral therapy based on WHO methods | Asia (WHO GHO)
🌏 525 observations · 15 Asia countries · 1990–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 525 observations of Estimated number of children needing antiretroviral therapy based on WHO methods data across 15 Asia countries, spanning 1990–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-estimated-number-of-children-needing-antiretroviral-therapy.
