datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.forgespectrum-114k
ForgeSpectrum (v3) — AI-Generated Image Detection with Reasoning Traces
ForgeSpectrum is a multi-domain corpus for AI-generated / manipulated image
detection, annotated by Gemini-2.5-Pro with structured forensic reasoning traces
(<fast>/<planning>/<reasoning>/<reflection>/<conclusion> patterns) plus per-image
attributes and suspicious-region notes.
v3 — what changed
v3 is the cleaned, balanced release:
3 domains: faces, scenes, id_cards (docs and scene_text… See the full description on the dataset page: https://huggingface.co/datasets/hardiksharma6555/forgespectrum-114k.celebamask-hq
Dataset Card for celebamask-hq
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/celebamask-hq")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/celebamask-hq.imagenet-hard-4K
Dataset Card for "Imagenet-Hard-4K"
Project Page - Paper - Github
ImageNet-Hard-4K is 4K version of the original ImageNet-Hard dataset, which is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard-4K.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.hard-intersection-multimodal-sample
Hard Intersection Multimodal Samples
Release Notes
Release
Description
v1.0.0
Initial public release.
v1.1.0
Added Unreal Engine assets.Fixed issues in the OpenDRIVE map data.Updated the README to improve documentation and usability.
Dataset Summary
Hard Intersection Multimodal Samples is a curated multimodal dataset of accident-prone urban intersection in Japan for autonomous driving research and development.It provides… See the full description on the dataset page: https://huggingface.co/datasets/dynamic-maps/hard-intersection-multimodal-sample.FairVision
Dataset Card: Harvard-FairVision
Dataset Summary
Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.harmful-contents
Harmful-Contents Dataset
A multi-label image dataset for harmful-content classification across eight PEGI-aligned categories.The dataset consists of 5,153 rights-cleared images, split into train/validation/test sets and annotated with both binary labels and mask fields for controlled negative sampling.
Dataset Structure
Harmful-Contents/
csv/
train.csv
val.csv
test.csv
data/
train/*.jpg
val/*.jpg
test/*.jpg
Each CSV contains:
name,
alcohol… See the full description on the dataset page: https://huggingface.co/datasets/onullusoy/harmful-contents.FairVLMed
Dataset Card: Harvard-FairVLMed
Dataset Summary
Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language.
This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.exp03-l23-hardening
Exp 03 / 03b — Klüver L2/3 hardening study (SDXL + SD 3.5)
Pre-registered study from the Operating System Hypothesis project.
Sweep classifier-free guidance across two architectures and score every output
blind, on two independent rubrics, for how far object structure has come apart.
The prediction was written down and committed before the run. The commit
dates in the GitHub repo are the proof.
What is here that is not on GitHub
The 860 generated PNGs. Every text… See the full description on the dataset page: https://huggingface.co/datasets/youssefhassan13/exp03-l23-hardening.PALL-VLM-data
PALL-VLM-data — Dental Vision-Language Dataset
The training dataset for Harisundar/PALL-VLM,
a dental vision-language model. It contains 32,884 records over 52,461 images,
formatted as image+text conversations for LLaVA-style instruction tuning.
Curated by: Harisundar R
Used by: Harisundar/PALL-VLM · PALL on GitHub
Language: English
Layout
vlm_train/
├── images/ # 52,461 dental images
├── train.jsonl # 29,667 records
├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.imagenet-hard
Dataset Card for "ImageNet-Hard"
Project Page - ArXiv - Paper - Github - Image Browser
Dataset Summary
ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their ability to… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard.MAVIS
MAVIS (Micro-surgical Artificial Vascular anastomosIS)
This dataset was presented in the paper: SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding.
Dataset Overview
MAVIS is a microsurgical dataset comprising 19 videos of artificial vascular anastomosis procedures performed by three expert microsurgeons at College of Medicine, Korea University, Republic of Korea.For each video frame, it provides:
Pixel-level… See the full description on the dataset page: https://huggingface.co/datasets/KIST-HARILAB/MAVIS.Fundus-CoT
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
glaucoma / not
train.jsonl
823
304 / 519
val.jsonl
92
46 / 46
test.jsonl
159
79 / 80
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train",
"final_diagnosis_GT":… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/Fundus-CoT.danbooru-tags-20260518Danbooru Dataset collected with my script.
Collected post ids: 1 ~ 11403815
Usage:
from datasets import load_dataset
dataset = load_dataset("u-haru/danbooru-tags-20260518", split="train")
ultrasound_images_diffusionThis dataset contains synthetic EDM2 generated images of ultrasound scan, as described in the paper A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets.
Code: https://github.com/xfetus/fetal-ultrasound-edm2
The class labels correspond to the following labels:
plane_classes = {
0: 'Other',
1: 'Maternal cervix',
2: 'Fetal abdomen',
3: 'Fetal brain',
4: 'Fetal femur',
5: 'Fetal thorax',
}
marvel-masterpieces-with-3dmesh
Dataset Card for reconstructions
Wait! Before you go, ❤️ the dataset! Let's get this trending!
This is a FiftyOne dataset with 255 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces-with-3dmesh.coco-cmfd
COCO-CMFD
A synthetic copy-move forgery dataset generated from MS-COCO 2017, with
source/target-separated ground truth for copy-move forgery detection
(CMFD).
Each sample takes one annotated COCO object, applies a mild affine
transform, and pastes it elsewhere in the same image at a location
that passes scene-plausibility checks (support surface, horizon band,
perspective scale, occlusion). Ground truth is provided as a 3-class
trimap, a binary mask, a 16 px patch-label grid… See the full description on the dataset page: https://huggingface.co/datasets/harshitajainn/coco-cmfd.FairGenMed
Dataset Card: FairGenMed
Dataset Summary
FairGenMed is the first dataset for studying fairness in medical generative models. It provides detailed quantitative clinical measurements alongside demographic annotations to investigate the semantic correlation between text prompts and anatomical regions across demographic subgroups. The dataset supports both generative model evaluation and downstream classification tasks for glaucoma detection.
This dataset accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairGenMed.visual_ai_at_neurips2025_jina_with_ocr
Dataset Card for harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr.index-cards-harvard-botany-metropolitan-flora
Card File of the Flora of the Metropolitan Parks (Harvard Botany Libraries, 1894–1895)
4,574 botanical specimen index cards from the Harvard University Botany
Libraries' Card File of the Flora of the Metropolitan Parks, 1894–1895 (bulk),
compiled by Walter Deane (1848–1930). Records flora of the Metropolitan Park
system around Boston — Middlesex Fells Reservation, Blue Hills, Norfolk County,
and adjacent areas — with one card per specimen entry: species, locality,
collection date… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-harvard-botany-metropolitan-flora.reLAIONet
Dataset Card: reLAIONet
Dataset Summary
reLAIONet is a manually proofread, web-sourced image classification benchmark aligned to ImageNet's 1,000-class label space. Sourced entirely from open web crawls (reLAION-400M) rather than Flickr, it provides a challenging out-of-distribution complement to ImageNet val and ImageNetV2 for evaluating discriminative and class-conditional generative models.
This dataset was introduced in: Fair Benchmarking of Emerging One-Step… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/reLAIONet.marvel-masterpieces
Dataset Card for marvel_masterpieces
Wait! Before you go, ❤️ the dataset! Let's get this trending!
This is a FiftyOne dataset with 255 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/marvel-masterpieces")
#… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces.cardd_workshop_post_03
Dataset Card for harpreetsahota/cardd_workshop_post_03
This is a FiftyOne dataset with 2816 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/cardd_workshop_post_03")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_03.cardd_workshop_post_01
Dataset Card for car_dd
This is a FiftyOne dataset with 2816 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/cardd_workshop_post_01")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_01.ssb_hard-ood
Dataset Card for SSB (hard) for OOD Detection
Dataset Details
Dataset Description
Original Dataset Authors: Sagar Vaze, Kai Han, Andrea Vedaldi, Andrew Zisserman
OOD Split Authors: Julian Bitterwolf, Maximilian Müller, Matthias Hein
Shared by: Eduardo Dadalto
License: unknown
Dataset Sources
Original Dataset Paper: http://arxiv.org/abs/2110.06207v2
First OOD Application Paper: http://arxiv.org/abs/2306.00826v1
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/detectors/ssb_hard-ood.ESC-10
Dataset Card for esc-10
This is a FiftyOne dataset with 400 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/ESC-10")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/ESC-10.isaac_on_images
Dataset Card for harpreetsahota/isaac_on_images
This is a FiftyOne dataset with 50 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/isaac_on_images")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/isaac_on_images.IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research… See the full description on the dataset page: https://huggingface.co/datasets/HarishBonu/IIIT-INDIC-HW-WORDS-Hindi.cardd_workshop_post_threed
Dataset Card for harpreetsahota/cardd_workshop_post_03
This is a FiftyOne dataset with 2816 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/cardd_workshop_post_threed")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_threed.
