datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
food101
Dataset Card for Food-101
Dataset Summary
This dataset consists of 101 food categories, with 101'000 images. For each class, 250 manually reviewed test images are provided as well as 750 training images. On purpose, the training images were not cleaned, and thus still contain some amount of noise. This comes mostly in the form of intense colors and sometimes wrong labels. All images were rescaled to have a maximum side length of 512 pixels.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ethz/food101.imagenet_1k_resized_256
Dataset Card for "imagenet_1k_resized_256"
Dataset summary
The same ImageNet dataset but all the smaller side resized to 256.
A lot of pretraining workflows contain resizing images to 256 and random cropping to 224x224, this is why 256 is chosen.
The resized dataset can also be downloaded much faster and consume less space than the original one.
See here for detailed readme.
Dataset Structure
Below is the example of one row of data. Note that the labels in… See the full description on the dataset page: https://huggingface.co/datasets/evanarlian/imagenet_1k_resized_256.eurosat
Dataset Card for EuroSAT
Dataset Source
Paper with code
Usage
from datasets import load_dataset
dataset = load_dataset('tranganke/eurosat')
Data Fields
The dataset contains the following fields:
image: An image in RGB format.
label: The label for the image, which is one of 10 classes:
0: annual crop land
1: forest
2: brushland or shrubland
3: highway or road
4: industrial buildings or commercial buildings
5: pasture land
6: permanent crop land… See the full description on the dataset page: https://huggingface.co/datasets/tanganke/eurosat.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.EuroSAT_RGB
EuroSAT RGB
EUROSAT RGB is the RGB version of the EUROSAT dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples.
Paper: https://arxiv.org/abs/1709.00029
Homepage: https://github.com/phelber/EuroSAT
Description
The EuroSAT dataset is a comprehensive land cover classification dataset that focuses on images taken by the ESA Sentinel-2 satellite. It contains a total of 27… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/EuroSAT_RGB.eurosat-rgb
EuroSat (RGB)
Description
A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/eurosat-rgb.Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.si_us_revolutionary_era_collections
Dataset Card for Smithsonian American Revolutionary Era Collections
Dataset Summary
A specially selected subset of the Smithsonian’s Open Access collections covering objects from 1770–1810 selected for the Revolution Crossroads project in honor of the 250th anniversary of the founding of the United States. Drawn from four museums—the National Museum of American History, National Postal Museum, Smithsonian American Art Museum, and National Portrait Gallery—the… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/si_us_revolutionary_era_collections.relaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.EpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.oxford-pets
Oxford-IIIT Pet Dataset
Images from The Oxford-IIIT Pet Dataset. Only images and labels have been pushed, segmentation annotations were ignored.
Homepage: https://www.robots.ox.ac.uk/~vgg/data/pets/
License:
Same as the original dataset.
mini_imagenet
Dataset Card for "mini_imagenet"
More Information needed
bone_marrow_cell_dataset
About This Dataset
Bone marrow biopsy is procedure applied to collect and examine bone marrow — the spongy tissue inside some of your larger bones.
This biopsy can show whether your bone marrow is healthy and making normal amounts of blood cells. Doctors use these procedures to diagnose and monitor blood and marrow diseases, cancers, as well as fevers of unknown origin.
The dataset contains a collection of over 170,000 de-identified, expert-annotated cells from the bone marrow… See the full description on the dataset page: https://huggingface.co/datasets/ekim15/bone_marrow_cell_dataset.oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.gz_evo
GZ Campaign Datasets
Dataset Summary
Galaxy Zoo volunteers label telescope images of galaxies according to their visible features: spiral arms, galaxy-galaxy collisions, and so on.
These datasets share the galaxy images and volunteer labels in a machine-learning-friendly format. We use these datasets to train our foundation models. We hope they'll help you too.
Curated by: Mike Walmsley
License: cc-by-nc-sa-4.0. We specifically require all models trained on these… See the full description on the dataset page: https://huggingface.co/datasets/mwalmsley/gz_evo.closed-open-eyes
👀 Open and Closed Eyes Dataset
Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟
📁 Dataset Structure
The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes.holisafe-bench
⚠️ CONTENT WARNING: This dataset contains potentially harmful and sensitive visual content including violence, hate speech, illegal activities, self-harm, sexual content, and other unsafe materials. Images are intended solely for safety research and evaluation purposes. Viewer discretion is strongly advised.
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model (CVPR'26 Findings)
🌐 Website | 📑 Paper
📋 HoliSafe-Bench Dataset… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/holisafe-bench.IntraOral_Gingivitis_Image_Captioning
A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING
Dataset Description
This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0.
This dataset contains 1,096 samples organized across multiple splits.
The dataset includes image data.
Splits
train: 732 samples
test: 182 samples
validation: 182 samples
Dataset Creation
This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.showui-web-processed
ShowUI-Web Processed
Flattened, normalized, and scenario-split version of showlab/ShowUI-web.
Each row is a single (instruction, UI element) pair with normalized bounding-box coordinates.
Schema
Column
Type
Description
sample_id
string
Unique row identifier ({row}_{element})
screenshot_id
string
Groups elements from the same screenshot
image_relpath
string
Relative path to the screenshot image
scenario
string
Website/domain inferred from the image path… See the full description on the dataset page: https://huggingface.co/datasets/e1879/showui-web-processed.NIH-Chest-XRay-Federated
NIH Chest X-ray Federated Learning Dataset
Federated learning splits designed for the [Cold Start:] Distributed AI Hack Berlin 2025.
The dataset is based on the NIH Chest X-ray14 dataset, which contains ~112,000 X-ray images from 30,805 unique patients, and models a federated learning scenario with non-IID characteristics across three hospitals, plus an out-of-distribution test set.
Dataset Description
The data was partitioned using a scoring algorithm that creates… See the full description on the dataset page: https://huggingface.co/datasets/exalsius/NIH-Chest-XRay-Federated.EMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.ego4d-random-views-20k
Ego4D Random Views Dataset
This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system.
Dataset Overview
Total Images: 20,000 high-quality frames
Image Format: PNG (1024×1024 resolution)
Source: Ego4D v2 dataset (52,665+ video files)
Sampling Method: Multi-process random sampling with maximum diversity
Generation Time: 797.57 seconds (~13 minutes)
Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.european_art
Dataset Card for DEArt: Dataset of European Art
Dataset Summary
DEArt is an object detection and pose classification dataset meant to be a reference for paintings between the XIIth and the XVIIIth centuries. It contains more than 15000 images, about 80% non-iconic, aligned with manual annotations for the bounding boxes identifying all instances of 69 classes as well as 12 possible poses for boxes identifying human-like objects. Of these, more than 50 classes are cultural… See the full description on the dataset page: https://huggingface.co/datasets/biglam/european_art.PathoROB-tolkach_esca
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-tolkach_esca.line-ex
LineEX
This dataset repo contains:
the uploaded train split with 397,993 images
the released test split with 20,000 images
Repo: 13point5/line-ex
Shared Schema
image_id
file_name
image
width
height
data_type
chart_elements
lines
chart_elements fields
annotation_id
category_id
category_name
bbox_xywh
area
text
line_id
lines fields
annotation_id
category_id
category_name
line_name
polyline_xy
raw_series_xy
area
Notes
The repo… See the full description on the dataset page: https://huggingface.co/datasets/13point5/line-ex.eurosat
EuroSAT-RGB Dataset
Dataset Description
The dataset comprises JPEG composite chips extracted from Sentinel-2 satellite imagery,
representing the Red, Green, and Blue bands.
It encompasses 27,000 labeled and geo-referenced images across 10 Land Use and Land Cover (LULC) classes
Dataset Structure
Splits : Train 80% Validation 10% Test 10%
(Kept the original dataset's label distribution consistent in each split)
Citation
Helber, P., Bischke, B.… See the full description on the dataset page: https://huggingface.co/datasets/cm93/eurosat.early_printed_books_font_detection
Early Printed Books Font Detection
Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper.
The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.fer2013-enhanced
FER2013 Enhanced: Advanced Facial Expression Recognition Dataset
The most comprehensive and quality-enhanced version of the famous FER2013 dataset for state-of-the-art emotion recognition research and applications.
🎯 Dataset Overview
FER2013 Enhanced is a significantly improved version of the landmark FER2013 facial expression recognition dataset. This enhanced version provides AI-powered quality assessment, balanced data splits, comprehensive metadata, and multi-format… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/fer2013-enhanced.radgenome-anatomy
RadGenome-Anatomy
RadGenome-Anatomy is a large-scale chest radiograph anatomy segmentation dataset
constructed from the RadGenome-ChestCT corpus
(originally based on CT-RATE).
It contains 25,692 volumetric studies (24,128 train / 1,564 validation), yielding paired
postero-anterior (PA) and lateral (LL) projection images at 384 × 384 resolution.
Across the two radiographic views, the dataset provides 10,790,646 fine-grained anatomy masks
over 210 canonical anatomy classes and 513… See the full description on the dataset page: https://huggingface.co/datasets/EvidenceAIResearch/radgenome-anatomy.
