datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zenodo-second-hand-fashion-v3
Second-Hand Fashion Dataset — wide (one row per garment)
Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row.
Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.child-emotion-drawings-pilot
Children's Emotional Drawings Pilot Dataset
A small balanced derived pilot dataset for experimental classification of emotional patterns in children's drawings.
Dataset
204 unique original drawings
492 total image instances
3 target classes: happiness, anxiety_depression, anger_aggression
Splits:
train: 432 images
validation: 30 images
test: 30 images
The dataset was created from the public anamelClassification dataset.
Original… See the full description on the dataset page: https://huggingface.co/datasets/stanislav-dykyi/child-emotion-drawings-pilot.AI-vs-Deepfake-vs-Real-Resized-Aug
🧠 AI vs Deepfake vs Real — Processed Version
This dataset is the result of preprocessing and augmentation applied to the original datasetprithivMLmods/AI-vs-Deepfake-vs-Real.
📘 Overview
This dataset contains a collection of images categorized into three main classes:
🟩 AI-generated
🟥 Deepfake
🟦 Real (authentic human faces)
It is designed for image classification tasks that aim to distinguish between AI-generated, deepfake, and real faces.… See the full description on the dataset page: https://huggingface.co/datasets/chintalaswathi/AI-vs-Deepfake-vs-Real-Resized-Aug.ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.synthetic-chest-xray-pneumonia
Synthetic Chest X-Ray Pneumonia Dataset
Dataset Description
This dataset contains synthetic chest X-ray images generated using Stable Diffusion 2.1
fine-tuned with DreamBooth on the hf-vision/chest-xray-pneumonia dataset.
Purpose
Created for a science fair project investigating whether synthetic medical images generated
by diffusion models can improve pneumonia classifier accuracy.
Research Question
Can synthetic chest X-ray images generated by a… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/synthetic-chest-xray-pneumonia.kawaii_chibi_avatar_dataset
Kawaii Chibi Avatar Dataset
This is the dataset used to train
Kawaii Chibi Avatar for Illustrious.
All images have a .txt file auto-tagged on Civitai.
All images were generated on SDXL using Kawaii Chibi Avatar for SDXL
License
License: CC BY 4.0
Attribution:
Kawaii Chibi Avatar Dataset © 2025 by Robb-0 is licensed under CC BY 4.0
COLD_chili_leaf_disease_classification
COLD Chili Leaf Disease Classification
A dataset for disease classification of chili leaves. The dataset contains raw and augmented versions.The raw dataset contains 532 images.Images per class:
cercospora: 152
healthy: 69
mites_and_trips: 107
nutritional: 102
powdery mildew: 102
The augmented dataset contains 10,974 images.Images per class:
cercospora: 2,217
healthy: 2,195
mites_and_trips: 2,504
nutritional: 2,029
powdery mildew: 2,029
This dataset is indexed on… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/COLD_chili_leaf_disease_classification.chitrak_leaf_disease_classification
Chitrak Leaf Disease Classification
A dataset for disease classification of Chitrak Leaves. The dataset contains 10,660 images across 3 classes: dried, healthy, and unhealthy.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{patil2024plumbago,
title={Plumbago Zeylanica (Chitrak) leaf image dataset: a comprehensive collection for botanical studies, herbal medicine research, and environmental… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/chitrak_leaf_disease_classification.CIFAR-10_Subset
CIFAR-10 — Subset
Stratified random subset of CIFAR-10.
Split
Rows
Per class
train
5,000
500
test
1,000
100
validation
500
50
Classes: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck
Images: 32 × 32 RGB | Seed: 42
Label Map
ID
Class
ID
Class
0
airplane
5
dog
1
automobile
6
frog
2
bird
7
horse
3
cat
8
ship
4deer
9
truck
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Chiranjeev007/CIFAR-10_Subset.MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/chinnawat2424/MM-Food-100K.Dior.Product.prices.China
Dior web scraped data
About the website
The luxury fashion industry in the Asia Pacific region, and in particular in China, is a rapidly expanding sector. As the middle class grows in wealth, there is an increasing demand for high-end goods from prestigious brands such as Christian Dior. The sector has been greatly influenced by technological innovations, the most notable of which is ecommerce. In recent years, there has been a significant shift in consumer behavior with… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Dior.Product.prices.China.Balenciaga.Product.prices.China
Balenciaga web scraped data
About the website
The fashion industry in the Asia Pacific region, particularly in China, is a hotbed of activity. It is one of the most lucrative markets in the world, spurred by a fast-growing middle class with an increased appetite for luxury products. The Chinese market, especially, plays host to many high-end, luxury fashion brands like Balenciaga. A significant transition has been noted in the mode of shopping, with a sharp turn towards… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Balenciaga.Product.prices.China.BinaryBreaKHisfootball-dataset_training
Dataset Description
This dataset contains football (soccer) field images captured from a tactical camera perspective. The dataset contains two images per frame;
One image highlights only the green color, while all other colors are grayed out.
The other image grays out the green color, while all other colors remain in full color.
The dataset is designed for computer vision research in sports analytics, including player tracking, field understanding, and tactical analysis.
Field… See the full description on the dataset page: https://huggingface.co/datasets/chimp-ll/football-dataset_training.Children_Drawings
Children's Drawings Dataset
This dataset contains children's drawings collected as part of a project conducted by the National Information Society Agency (NIA) of Korea.
Dataset Structure
The dataset includes:
Drawings of houses, trees, male figures, and female figures
Both raw source images and labeled data
Training and validation splits
Categories
집 (House)
나무 (Tree)
남자사람 (Male figure)
여자사람 (Female figure)
File Naming Convention
Files are named… See the full description on the dataset page: https://huggingface.co/datasets/ironDong/Children_Drawings.STL-10_Subset
STL-10 — Subset
Stratified random subset of STL-10.
Split
Rows
Per class
train
5,000
500
test
1,000
100
validation
500
50
Classes: airplane, bird, car, cat, deer, dog, horse, monkey, ship, truck
Images: 96 × 96 RGB | Seed: 42
from datasets import load_dataset
ds = load_dataset("Chiranjeev007/STL-10_Subset")
sample = ds["train"][0]
sample["image"] # PIL Image 96×96 RGB
sample["label"] # int 0–9
Hermes.Product.prices.China
Hermes web scraped data
About the website
The luxury goods industry in the Asia Pacific, particularly in China, is experiencing significant growth driven by rising wealth and changing consumer preferences. Hermes, a high-end luxury brand, is a notable player within this affluent sector. This industrys success is tied to the robust ecommerce landscape in China, characterized by innovative digital platforms with extensive customer reach. The dataset examined contains… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Hermes.Product.prices.China.Loro.Piana.Product.prices.China
Loro Piana web scraped data
About the website
The luxury fashion industry is an influential sector in the Asia Pacific, particularly in China, which exhibits an increasing influence on the global landscape. It is driven by rising consumer demand for high-end fashion and luxury goods. With the advent of technology, the e-commerce platform has become a new playground for luxury brands like Loro Piana. E-commerce in China provides an opportunity for these brands to extend… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Loro.Piana.Product.prices.China.Burberry.Product.prices.China
Burberry web scraped data
About the website
The luxury fashion industry in the Asia Pacific region, particularly in China, has seen a significant shift towards digitalization. Online shopping, fuelled by the growth of Ecommerce, has become a major sales channel for high-end labels like Burberry. This growth in online sales has outpaced that of the offline sector, making e-commerce a key driver for the luxury fashion sector. Chinese consumption of luxury goods is turning… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Burberry.Product.prices.China.SVHN-10_Subset
SVHN (Street View House Numbers) — Subset 🔢
Stratified random subset of SVHN,
sourced from ufldl-stanford/svhn.
Splits
Split
Rows
Classes
Per class
train
5,000
10
~500
test
1,000
10
~100
validation
500
10
~50
About SVHN
The Street View House Numbers (SVHN) dataset is a real-world digit recognition dataset
obtained from Google Street View images.
Full train set: 73,257 images
Full test set: 26,032 images
Extra set: 531,131 images (not… See the full description on the dataset page: https://huggingface.co/datasets/Chiranjeev007/SVHN-10_Subset.ppi-chine-ccnu-01
Video Dataset - ccnu-01
Dataset Description
This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips.
Dataset Structure
frames/ — extracted frames (first frame from each segment)
segments/ — video clips for each annotation interval
annotations/ — original JSON annotation
transcriptions/ — transcription files (full_transcription.txt + per segment)
dataset.csv — mapping… See the full description on the dataset page: https://huggingface.co/datasets/icomgpu/ppi-chine-ccnu-01.cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/Chilli2812/cifar10.
