datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChildGaitDecoding Children's Gait Behavior
ECCV 2026
Yifan Shen1,2,*,
Boyi Li1,*,
Meihuan Huang2,3,4,*,
Yuanzhe Liu1,*,
Xu Cao1,2,*,§,
Jinyang Jin1,
Zhengyuan Li1,
Anglin Liu5,
Junho Kim1,
Jingyuan Zhu2,
Fangzhou Lan2,
Jianguo Cao2,3,
Jintai Chen5,
Ismini Lourentzou1,
James M. Rehg1,†
1 University of Illinois Urbana-Champaign
2 PediaMed AI
3 Shenzhen Children's Hospital
4 Hong Kong Polytechnic University… See the full description on the dataset page: https://huggingface.co/datasets/PediaMedAI/ChildGait.the-un-laion-templeAll files uploaded. Enjoy!
Dataset Card for The Unlaion Temple
Dataset Details
Dataset Description
Laion-5B is still not public, so we decided to create our own dataset.
The Unlaion Temple is a raw dataset of CommonCrawl images (Estimated to be a total of 2 Billion urls). We haven't verified whether the links in this dataset are functional.
You are responsible for handling the data.
We've made some improvements to the dataset based on user feedback:
All… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Chiharu/the-un-laion-temple.zenodo-second-hand-fashion-v3
Second-Hand Fashion Dataset — wide (one row per garment)
Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row.
Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.ChildGaitDecoding Children's Gait Behavior
ECCV 2026
Yifan Shen1,2,*,
Boyi Li1,*,
Meihuan Huang2,3,4,*,
Yuanzhe Liu1,*,
Xu Cao1,2,*,§,
Jinyang Jin1,
Zhengyuan Li1,
Anglin Liu5,
Junho Kim1,
Jingyuan Zhu2,
Fangzhou Lan2,
Jianguo Cao2,3,
Jintai Chen5,
Ismini Lourentzou1,
James M. Rehg1,†
1 University of Illinois Urbana-Champaign
2 PediaMed AI
3 Shenzhen Children's Hospital
4 Hong Kong Polytechnic University… See the full description on the dataset page: https://huggingface.co/datasets/gopichand143/ChildGait.child-emotion-drawings-pilot
Children's Emotional Drawings Pilot Dataset
A small balanced derived pilot dataset for experimental classification of emotional patterns in children's drawings.
Dataset
204 unique original drawings
492 total image instances
3 target classes: happiness, anxiety_depression, anger_aggression
Splits:
train: 432 images
validation: 30 images
test: 30 images
The dataset was created from the public anamelClassification dataset.
Original… See the full description on the dataset page: https://huggingface.co/datasets/stanislav-dykyi/child-emotion-drawings-pilot.galaxy-chirality-catalog
DESI Legacy Galaxy Chirality Catalog
This dataset accompanies the current Paper IV manuscript, An Observed-Label Chirality-Dipole Null in 949,584 High-Confidence DESI Spirals and an 8.5-Million-Galaxy Catalog.
The primary high-confidence observed-label statistic is consistent with zero under fixed-occupancy label randomization (z=0.7053169638, one-sided empirical-rank p=0.2246775322). This is not a calibrated true-spin, physical-amplitude, or primordial-parity bound.… See the full description on the dataset page: https://huggingface.co/datasets/bamfai/galaxy-chirality-catalog.vqa-test-studies
UCSF-PDGM Test Subset (VQA Demo)
A three-study subset of the UCSF Preoperative Diffuse Glioma MRI (UCSF-PDGM) dataset, mirrored here as test fixtures for a web-based visual question answering (VQA) demo. The full source dataset is publicly available on The Cancer Imaging Archive (TCIA).
Contents
Three preoperative brain MRI studies from patients with diffuse glioma:
UCSF-PDGM-0159_nifti/
UCSF-PDGM-0338_nifti/
UCSF-PDGM-0529_nifti/
Each study folder contains… See the full description on the dataset page: https://huggingface.co/datasets/chihhua0908/vqa-test-studies.AI-vs-Deepfake-vs-Real-Resized-Aug
🧠 AI vs Deepfake vs Real — Processed Version
This dataset is the result of preprocessing and augmentation applied to the original datasetprithivMLmods/AI-vs-Deepfake-vs-Real.
📘 Overview
This dataset contains a collection of images categorized into three main classes:
🟩 AI-generated
🟥 Deepfake
🟦 Real (authentic human faces)
It is designed for image classification tasks that aim to distinguish between AI-generated, deepfake, and real faces.… See the full description on the dataset page: https://huggingface.co/datasets/chintalaswathi/AI-vs-Deepfake-vs-Real-Resized-Aug.ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.PASD
PASD — Placenta Accreta Spectrum MRI Dataset
A 3D MRI dataset for Placenta Accreta Spectrum (PAS) diagnosis with
voxel-level lesion masks and case-level diagnostic labels. This dataset
accompanies the paper:
3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum, IEEE Transactions on Image Processing.
Source code for the proposed 3DSAMba method:
https://github.com/Drchip61/PASD.
Dataset Summary
Split
Cases
Negative (label=0)… See the full description on the dataset page: https://huggingface.co/datasets/ChipYTY/PASD.ImageNet_FID50K
ImageNet-256 Inception pool3 features (50k) — reference and VAR samples
Precomputed Inception-v3 pool3 features (2048-d) for a 50,000-image ImageNet-256 reference set and
for 50,000 generated samples from each of the three VAR
checkpoints (Tian et al., NeurIPS 2024).
These are the exact features behind the energy-distance / FID comparison in our work. They let you
reproduce distributional metrics without re-running Inception over 200,000 images (~1 GPU-hour),
and without… See the full description on the dataset page: https://huggingface.co/datasets/chicagoypark/ImageNet_FID50K.synthetic-chest-xray-pneumonia
Synthetic Chest X-Ray Pneumonia Dataset
Dataset Description
This dataset contains synthetic chest X-ray images generated using Stable Diffusion 2.1
fine-tuned with DreamBooth on the hf-vision/chest-xray-pneumonia dataset.
Purpose
Created for a science fair project investigating whether synthetic medical images generated
by diffusion models can improve pneumonia classifier accuracy.
Research Question
Can synthetic chest X-ray images generated by a… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/synthetic-chest-xray-pneumonia.kawaii_chibi_avatar_dataset
Kawaii Chibi Avatar Dataset
This is the dataset used to train
Kawaii Chibi Avatar for Illustrious.
All images have a .txt file auto-tagged on Civitai.
All images were generated on SDXL using Kawaii Chibi Avatar for SDXL
License
License: CC BY 4.0
Attribution:
Kawaii Chibi Avatar Dataset © 2025 by Robb-0 is licensed under CC BY 4.0
phantom-ledger-chimera
Phantom Ledger: Chimera v2 — Synthetic Multi-Modal Fraud Ledger
Synthetic dataset (MIT) generated from scratch via dataset/raw/generate_chimera_v2.py (seed 42). No real PII.
18,216 transactions (4.54% fraud, train 4.34% / test 5.44%, drift 2024-09-01)
1,200 merchants (KB) + 2,500 users + 18,216 receipt images (420×220)
Modalities: ledger.csv (28 cols), receipts/*.png (border leak + tampered amount), merchant_kb.jsonl (stale PRE), user_sequences.jsonl (future-leak)
See… See the full description on the dataset page: https://huggingface.co/datasets/akiii1234/phantom-ledger-chimera.PseudoKitchens
Dataset Card for PseudoKitchens
Dataset Summary
PseudoKitchens is a synthetic dataset of photorealistic 3D kitchen renders with ground-truth concept (object) annotations. It is designed for tasks such as recipe classification and spatial concept localisation. The dataset consists of kitchen scenes containing ingredients for recipe classification tasks, with each scene accompanied by annotations that describe the location of every ingredient in it.
PseudoKitchens-2 is a… See the full description on the dataset page: https://huggingface.co/datasets/yoda-chicken/PseudoKitchens.COLD_chili_leaf_disease_classification
COLD Chili Leaf Disease Classification
A dataset for disease classification of chili leaves. The dataset contains raw and augmented versions.The raw dataset contains 532 images.Images per class:
cercospora: 152
healthy: 69
mites_and_trips: 107
nutritional: 102
powdery mildew: 102
The augmented dataset contains 10,974 images.Images per class:
cercospora: 2,217
healthy: 2,195
mites_and_trips: 2,504
nutritional: 2,029
powdery mildew: 2,029
This dataset is indexed on… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/COLD_chili_leaf_disease_classification.chitrak_leaf_disease_classification
Chitrak Leaf Disease Classification
A dataset for disease classification of Chitrak Leaves. The dataset contains 10,660 images across 3 classes: dried, healthy, and unhealthy.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{patil2024plumbago,
title={Plumbago Zeylanica (Chitrak) leaf image dataset: a comprehensive collection for botanical studies, herbal medicine research, and environmental… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/chitrak_leaf_disease_classification.chinese-mainland-landmark-images
Chinese Mainland Landmark Reference Images Dataset
数据集描述
本数据集包含中国地标建筑的参考图像。
数据集结构
images/: 图像文件目录
metadata.jsonl: 元数据文件(JSONL 格式)
元数据字段
name: 地标建筑名称
adcode: 行政区划代码(6位)
city: 所在城市和区县(格式:城市名·区名)
image_file: 图像文件名
数据集统计
总图像数: 7911
图像格式: PNG
使用示例
from datasets import load_dataset
dataset = load_dataset("86Cao/chinese-mainland-landmark-images")
CIFAR-10_Subset
CIFAR-10 — Subset
Stratified random subset of CIFAR-10.
Split
Rows
Per class
train
5,000
500
test
1,000
100
validation
500
50
Classes: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck
Images: 32 × 32 RGB | Seed: 42
Label Map
ID
Class
ID
Class
0
airplane
5
dog
1
automobile
6
frog
2
bird
7
horse
3
cat
8
ship
4deer
9
truck
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Chiranjeev007/CIFAR-10_Subset.MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/chinnawat2424/MM-Food-100K.Dior.Product.prices.China
Dior web scraped data
About the website
The luxury fashion industry in the Asia Pacific region, and in particular in China, is a rapidly expanding sector. As the middle class grows in wealth, there is an increasing demand for high-end goods from prestigious brands such as Christian Dior. The sector has been greatly influenced by technological innovations, the most notable of which is ecommerce. In recent years, there has been a significant shift in consumer behavior with… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Dior.Product.prices.China.chineseBalenciaga.Product.prices.China
Balenciaga web scraped data
About the website
The fashion industry in the Asia Pacific region, particularly in China, is a hotbed of activity. It is one of the most lucrative markets in the world, spurred by a fast-growing middle class with an increased appetite for luxury products. The Chinese market, especially, plays host to many high-end, luxury fashion brands like Balenciaga. A significant transition has been noted in the mode of shopping, with a sharp turn towards… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Balenciaga.Product.prices.China.BinaryBreaKHisfootball-dataset_training
Dataset Description
This dataset contains football (soccer) field images captured from a tactical camera perspective. The dataset contains two images per frame;
One image highlights only the green color, while all other colors are grayed out.
The other image grays out the green color, while all other colors remain in full color.
The dataset is designed for computer vision research in sports analytics, including player tracking, field understanding, and tactical analysis.
Field… See the full description on the dataset page: https://huggingface.co/datasets/chimp-ll/football-dataset_training.Children_Drawings
Children's Drawings Dataset
This dataset contains children's drawings collected as part of a project conducted by the National Information Society Agency (NIA) of Korea.
Dataset Structure
The dataset includes:
Drawings of houses, trees, male figures, and female figures
Both raw source images and labeled data
Training and validation splits
Categories
집 (House)
나무 (Tree)
남자사람 (Male figure)
여자사람 (Female figure)
File Naming Convention
Files are named… See the full description on the dataset page: https://huggingface.co/datasets/ironDong/Children_Drawings.STL-10_Subset
STL-10 — Subset
Stratified random subset of STL-10.
Split
Rows
Per class
train
5,000
500
test
1,000
100
validation
500
50
Classes: airplane, bird, car, cat, deer, dog, horse, monkey, ship, truck
Images: 96 × 96 RGB | Seed: 42
from datasets import load_dataset
ds = load_dataset("Chiranjeev007/STL-10_Subset")
sample = ds["train"][0]
sample["image"] # PIL Image 96×96 RGB
sample["label"] # int 0–9
Parent-Child-Outfit-Image-Classification-Dataset
Parent-Child Outfit Image Classification Dataset
With the rapid development of retail e-commerce, the parent-child outfit market is also gradually expanding, and consumer demand for parent-child outfit products is increasing. However, existing product classification systems often fail to accurately identify and classify parent-child outfits, leading to poor user experience and affecting sales performance. Current solutions lack in labeling accuracy and data integrity, making it… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Parent-Child-Outfit-Image-Classification-Dataset.Hermes.Product.prices.China
Hermes web scraped data
About the website
The luxury goods industry in the Asia Pacific, particularly in China, is experiencing significant growth driven by rising wealth and changing consumer preferences. Hermes, a high-end luxury brand, is a notable player within this affluent sector. This industrys success is tied to the robust ecommerce landscape in China, characterized by innovative digital platforms with extensive customer reach. The dataset examined contains… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Hermes.Product.prices.China.Loro.Piana.Product.prices.China
Loro Piana web scraped data
About the website
The luxury fashion industry is an influential sector in the Asia Pacific, particularly in China, which exhibits an increasing influence on the global landscape. It is driven by rising consumer demand for high-end fashion and luxury goods. With the advent of technology, the e-commerce platform has become a new playground for luxury brands like Loro Piana. E-commerce in China provides an opportunity for these brands to extend… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Loro.Piana.Product.prices.China.
