datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Brain-Tumour-MRI
Dataset Card for Brain Tumour MRI dataset
A collection of Brain scans covering three different types of tumours and as well as a control class.
Dataset Details
Dataset Description
The Dataset contains ~7000 MRI scans of the brain corresponding to 4 classes: glioma, meningioma, notumor & pituitary.
The dataset has already been split into train/test sets.
Dataset Creation
Source
This dataset was compiled and uploaded to Kaggle by Masoud… See the full description on the dataset page: https://huggingface.co/datasets/Kaynaaf/Brain-Tumour-MRI.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.DisasterM3
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
Junjue Wang*,
Weihao Xuan*,
Heli Qi, Zhihao Liu, Kunyi Liu, Yuhan Wu, Hongruixuan Chen,
Jian Song
Junshi Xia, Zhuo Zheng, Naoto Yokoya†
* Equal Contributions
† Corresponding Author
Paper: https://arxiv.org/abs/2505.21089
Code: https://github.com/Junjue-Wang/DisasterM3
Highlights
DisasterM3 includes 26,988 bi-temporal satellite images and 123k instruction pairs across… See the full description on the dataset page: https://huggingface.co/datasets/Kingdrone-Junjue/DisasterM3.kmnist
KMNIST Dataset
lassify images from the KMNIST dataset into one of the 10 classes, representing different Japanese characters.
data-csgo-weapon-classification
Dataset for project: csgo-weapon-classification
Dataset Description
This dataset has for project csgo-weapon-classification was collected with the help of a bulk google image downloader.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<1768x718 RGB PIL image>",
"target": 0
},
{
"image": "<716x375 RGBA PIL image>"… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-csgo-weapon-classification.ArtiBench
ArtiBench: Artifact Detection Benchmark
Dataset Structure
Artifact-positive samples:
{
"id": "3qotz3zm",
"has_artifacts": true,
"explanation": "The image presents an aerial view of downtown Manhattan with an unusual twist. A large Ferris wheel, reminiscent of the Millennium Wheel, is oddly positioned next to the skyscrapers, appearing to be fused with the buildings below. ...",
"bboxes": [[114, 253, 432, 694]]
}
Artifact-negative samples:
{
"id": "nkzk0lqs"… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/ArtiBench.amfitrite-inland-waters-hab-sentinel2
Dataset Card for Amfitrite-Inland-Waters-HAB-Sentinel2
This dataset contains multispectral Sentinel-2 satellite imagery tiles focused on inland water bodies, classified by the severity of Harmful Algal Blooms (HABs).
It is designed to train Deep Learning models (like CNNs) for environmental monitoring.
Dataset Details
Dataset Description
Amfitrite-Inland-Waters-HAB-Sentinel2 is a specialized dataset designed for the detection and classification of… See the full description on the dataset page: https://huggingface.co/datasets/kostaspic/amfitrite-inland-waters-hab-sentinel2.oita-ken-strawberries
oita-ken-strawberries
This dataset contains 5,000 images of ベリーツ (Beriitsu), a premium strawberry variety grown in Oita Prefecture, Japan.
Images in dataset/input/* are organized by harvest year and grading category.
The dataset/input/**/include directory contains cleaned strawberry images.
Data cleansing was performed using background_erase.
Conversion to Pascal VOC
dataset/input/input.json and dataset/input/** can be processed with image_data_augmentation… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/oita-ken-strawberries.Breast_Pathology_Image_JPG
BreastPathDB
BreastPathDB is a breast cancer histopathology dataset for whole-slide image based subtype classification. It contains 770 de-identified whole-slide images, including 165 H&E slides and 605 IHC slides, with clinical subtype labels, curated metadata, and patient-level train/validation/test splits.
The dataset is intended for computational pathology benchmarking, weakly supervised WSI classification, and cross-cohort method development.
Browse files and metadata… See the full description on the dataset page: https://huggingface.co/datasets/kmmuleelab/Breast_Pathology_Image_JPG.plant_image_degradation
Plant Disease Degradation 224
Dataset Summary
This dataset contains 224×224 plant disease images and their synthetically degraded variants for image-classification and robustness experiments. It is designed for transfer-learning workflows that require a fixed input size of 224×224 while still exposing models to realistic quality issues that occur in field capture.
The degraded images are generated from clean plant-leaf images using a field-oriented degradation… See the full description on the dataset page: https://huggingface.co/datasets/kunley2/plant_image_degradation.nut_defect_detection
Nut Defect Classification
Synthetic Industrial Quality Inspection Dataset
Nut Defect Classification (Synthetic Dataset)
This dataset is a synthetic collection of industrial nut images designed for image classification tasks, specifically focusing on defect detection in manufacturing pipelines. It serves as a benchmark and training resource for computer vision algorithms used in quality assurance.… See the full description on the dataset page: https://huggingface.co/datasets/Kinzaaa/nut_defect_detection.chest-xray-classification
Dataset Labels
['NORMAL', 'PNEUMONIA']
Number of Images
{'train': 4077, 'test': 582, 'valid': 1165}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("keremberke/chest-xray-classification", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/mohamed-traore-2ekkp/chest-x-rays-qjmia/dataset/2
Citation… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/chest-xray-classification.PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.kitchen_utensils_5k
Kitchen Utensils Dataset (5K Images)
Dataset Description
This dataset contains 5,000 images of kitchen utensils for object detection/classification.
Classes (10 total):
bowl
fork
glass
knife
mug
pan
plate
spatula
spoon
whisk
Dataset Structure
Train: ~3,500 images
Validation: ~1,000 images
Test: ~500 images
File Format
Images: JPG format
Annotations: CSV file with binary labels for each class
metadata.csv: Combined metadata for all… See the full description on the dataset page: https://huggingface.co/datasets/rypow/kitchen_utensils_5k.sonata-dental-dataset
Sonata Dental Dataset
牙科疾病检测与分类数据集(基于 Sonata 采集数据整理)。
⚠️ 原始完整版(含 unknown 类)已备份,未随本仓库发布。
目录结构
Sonata/
├── image/ # 检测图像(1645 张)
├── label/ # 检测标注(labelme JSON,1645 个,与 image 一一对应)
└── Periodontal_Disease/ # 牙周病分类子集(按类别分目录)
├── gingival_diseases/ # 1103 张
├── non_periodontal_disease/ # 591 张
└── periodontitis/ # 661 张
1. 检测子集(Detection)
图像:image/,1645… See the full description on the dataset page: https://huggingface.co/datasets/Kellection/sonata-dental-dataset.kriyam-tamperflow
Kriyam TamperFlow
The first document tampering detection benchmark built specifically for Indian documents, with a built-in compression stress-test that exposes how quickly forensic models degrade on real-world scanned material.
Dataset Summary
State-of-the-art document forgery detectors — CAT-Net, DTD, MVSS-Net, CAFTB, and others — rely on JPEG compression artifacts as their primary forensic signal: inconsistencies in DCT coefficients, block boundaries, and… See the full description on the dataset page: https://huggingface.co/datasets/kriyam-ai/kriyam-tamperflow.coyo-labeled-300m
Dataset Card for COYO-Labeled-300M
Dataset Summary
COYO-Labeled-300M is a dataset of machine-labeled 300M images-multi-label pairs. We labeled subset of COYO-700M with a large model (efficientnetv2-xl) trained on imagenet-21k. We followed the same evaluation pipeline as in efficientnet-v2. The labels are top 50 most likely labels out of 21,841 classes from imagenet-21k. The label probabilies are provided rather than label so that the user can select threshold of their… See the full description on the dataset page: https://huggingface.co/datasets/kakaobrain/coyo-labeled-300m.data-food-classification
Dataset for project: food-classification
Dataset Description
This dataset has been processed for project food-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<308x512 RGB PIL image>",
"target": 0
},
{
"image": "<512x512 RGB PIL image>",
"target": 0
}]
Dataset Fields
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-food-classification.Mars-Analog-Dunes-10
Mars-Analog-Dunes-10 Dataset
Dataset Description
Mars-Analog-Dunes-10 is a curated remote sensing dataset designed for Earth-Mars Comparative Planetology. It contains high-quality satellite imagery of Earth's sand dunes that serve as morphological analogs to features found on Mars.
Key Features
Mars Analog Focus: Selected from Earth regions (e.g., deserts in China, Africa) known for their similarity to Martian geomorphology.
10 Fine-Grained Classes:… See the full description on the dataset page: https://huggingface.co/datasets/Keiyoo/Mars-Analog-Dunes-10.NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset
Dataset Card for NOAA-ESD-CORAL-Bleaching Classification Dataset v1
Overview
For the development of machine learning models to classify coral health, specifically identifying healthy hard coral (CORAL) and bleached hard coral (CORAL_BL).This dataset contains underwater imagery collected by NOAA's Ecosystem Sciences Division (ESD) and other benthic surveys.
Labels
Label
Name
Functional Group
CORAL
Healthy Hard Coral
Hard Coral
CORAL_BL… See the full description on the dataset page: https://huggingface.co/datasets/Kshoarya-8/NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset.NREL_Sky_Imagery
NREL SRRL Minute-Resolution Sky Imagery Dataset
Homepage
https://huggingface.co/datasets/knl2366/NREL_Sky_Imagery
Paper
Hammond & Korgel (2026), Journal of Data-centric Machine Learning Research
Contact
Joshua E. Hammond (jeh5975@utexas.edu)
Summary
A continuously growing dataset of minute-resolution sky images from the EKO ASI-16 all-sky imager at NREL's Solar Radiation Research Laboratory (SRRL) in Golden, Colorado (39.742°N, 105.180°W, 1829… See the full description on the dataset page: https://huggingface.co/datasets/knl2366/NREL_Sky_Imagery.ffhq-image-attribution
FFHQ Image Attribution
A public benchmark for FFHQ model attribution, built from twelve face generators spanning GAN, VAE, and diffusion families.
Version v2 includes 10,000 images from each of 12 FFHQ-trained generators for a total of 120,000 images.
Why this dataset
Same image domain across multiple FFHQ generators makes source attribution cleaner and easier to study.
Public metadata links each image to its source model, family, release, seed, and file integrity… See the full description on the dataset page: https://huggingface.co/datasets/kaikaiyao/ffhq-image-attribution.old-doors-50-8-commercial
European Architectural Doors Dataset — 50 Authentic Details
📋 Description
Curated collection of 50 high-resolution photographs documenting authentic European architectural doors — from rustic Italian farmhouses and Tuscan alleyways to ornate stone archways, Gothic cathedral entrances, medieval iron gates, historic French academy portals, and charming countryside barn doors.
Each image includes comprehensive CSV metadata with 13 classification fields optimized for… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/old-doors-50-8-commercial.food-category-classification-v2.0
Dataset for project: food-category-classification-v2.0
Dataset Description
This dataset for project food-category-classification-v2.0 was scraped with the help of a bulk google image downloader.
Dataset Structure
Dataset Fields
The dataset has the following fields (also called "features"):
{
"image": "Image(decode=True, id=None)",
"target": "ClassLabel(names=['Bread', 'Dairy', 'Dessert', 'Egg', 'Fried Food', 'Fruit', 'Meat', 'Noodles', 'Rice'… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/food-category-classification-v2.0.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.klingai-images
KlingAI Image
This dataset contains 3791 image samples extracted from the nyuuzyou/klingai dataset.
Dataset Structure
Media files: image files in images/ directory
Metadata: image_metadata.parquet contains all metadata including:
Original resource URLs
Dimensions (width, height)
Duration (for videos)
All other fields from the source dataset
Source
Original dataset: nyuuzyou/klingai
Processed and uploaded by the bitmind team for use in media generation… See the full description on the dataset page: https://huggingface.co/datasets/bitmind/klingai-images.CraftSight-Minecraft
CraftSight
Frame-level multi-label visual annotations for Minecraft agents.
Dataset creation toolkit: CraftSight is built and maintained with CraftSight Labeler, an open-source, browser-based annotation tool and reproducible release pipeline for multi-label Minecraft vision datasets. It supports manual and model-assisted labeling, structured game-state annotations, and trajectory-safe train/validation/test splits.
CraftSight provides Minecraft gameplay frames annotated… See the full description on the dataset page: https://huggingface.co/datasets/Krows7/CraftSight-Minecraft.kids-and-teens-selfie-dataset
Age Estimation
The dataset consists of 6,000 high-quality facial images from 300 people (children and teenagers), featuring a diverse range of facial features, poses, and attributes. It is designed for research and development in age estimation, facial recognition for younger demographics, and understanding media use patterns on social media platforms.
By utilizing this dataset, researchers and developers can advance their models for responsible technology, ensuring safer… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/kids-and-teens-selfie-dataset.1-engineering-repair-50-commercial
Industrial Mechanical & Electrical Components Dataset — 50 Authentic Details
📋 Description
Curated collection of 50 high-resolution photographs documenting authentic industrial mechanical and electrical components — from electric motors and solenoid valves to circuit boards, wiring harnesses, ball bearings, gearboxes, caster wheels, and battery packs.
Each image includes comprehensive CSV metadata with 13 classification fields optimized for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/1-engineering-repair-50-commercial.data-food-category-classification
Dataset for project: food-category-classification
Dataset Description
This dataset is for project food-category-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<512x512 RGB PIL image>",
"target": 0
},
{
"image": "<512x512 RGB PIL image>",
"target": 0
}]
Dataset Fields
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-food-category-classification.
