datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hyperspectral-orchard
Living Optics Orchard Dataset
Overview
This dataset contains 435 images of captured in one of the UK's largest orchards, using the Living Optics Camera.
The data consists of RGB images, sparse spectral samples and instance segmentation masks.
The dataset is derived from 44 unique raw files corresponding to 435 frames.
Therefore, multiple frames could originate from the same raw file.
This structure emphasized the need for a split strategy that avoided data leakage.
To… See the full description on the dataset page: https://huggingface.co/datasets/LivingOptics/hyperspectral-orchard.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.American-Sign-Language-MNIST
Dataset Card for ASL-MNIST
This is a FiftyOne dataset with 34,627 samples of American Sign Language (ASL) alphabet images, converted from the original Kaggle Sign Language MNIST dataset into a format optimized for computer vision workflows.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/American-Sign-Language-MNIST.military_vehicles
Citation
If you use this dataset, please cite the following paper:
@article{kricheli2024error,
title={Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge},
author={Kricheli, Joshua Shay and Vo, Khoa and Datta, Aniruddha and Ozgur, Spencer and Shakarian, Paulo},
journal={arXiv preprint arXiv:2407.15192},
year={2024}
}
dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/LiamLian0727/UIIS.DiTFakeHere is the released dataset (DiTFake) for Synthetic Image Detection (SID) proposed in our paper.
Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective
This dataset contains 30,000 images in total, including synthetic images generated by three recent DiT-based models (Flux, PixArt, and SD3) and equal numbers of real images from COCO.
More implementation details can be found in our GitHub repository.
Android-in-the-Wild
Android in the Wild (AITW)
This is a mirror of Google's Android in the Wild (AITW) dataset, re-hosted on Hugging Face for easier community access.
Original Source
Paper: Android in the Wild: A Large-Scale Dataset for Android Device Control
Original Repository: google-research/google-research/tree/master/android_in_the_wild
Dataset Description
Android in the Wild (AITW) is a large-scale dataset for Android device control. It contains human demonstrations of… See the full description on the dataset page: https://huggingface.co/datasets/leosltl/Android-in-the-Wild.mmcows
MmCows: A Multimodal Dataset for Dairy Cattle Monitoring
Details of the dataset and benchmarks are available here.
For a quick overview of the dataset, please check this video.
Instruction for downloading
1. Install requirements
pip install huggingface_hub
See the file structure here for the next step.
2. Download a file individually
To download visual_data.zip to your local-dir, use command line:
huggingface-cli download \
neis-lab/mmcows \… See the full description on the dataset page: https://huggingface.co/datasets/neis-lab/mmcows.laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.EEG_Image_decode
EEG Image Decode — Dataset and Checkpoints
This dataset accompanies the NeurIPS 2024 paper:
Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion Dongyang Li · Chen Wei · Shiying Li · Jiachen Zou · Quanying Liu
It packages the preprocessed EEG recordings, stimulus-image visual features, VAE latent codes, trained EEG embeddings, fine-tuned checkpoints, and generated images needed to reproduce both the image retrieval and image reconstruction experiments.… See the full description on the dataset page: https://huggingface.co/datasets/LidongYang/EEG_Image_decode.SciPanelForge
SciPanelForge
SciPanelForge is a paired scientific panel-to-code dataset prepared from the
audited clean benchmark pool of SciFigure2Code/benchmark_ready. Each sample contains an original
paper panel, a refined reproduction result, and the Python code used for the
refined reproduction.
Project Links
GitHub: https://github.com/littlepeachs/NaturePanelForge
Website: https://uu543493-83c1-74a94416.nma1.seetacloud.com:8448/
Dataset Contents
Samples:… See the full description on the dataset page: https://huggingface.co/datasets/littlepeachs/SciPanelForge.LUMA
LUMA
A Benchmark Dataset for Learning from Uncertain and Multimodal Data
📄
📷
🎵
📊
❓
Multimodal Uncertainty Quantification at Your Fingertips
The LUMA dataset is a multimodal dataset, including audio, text, and image modalities, intended for benchmarking multimodal learning and multimodal uncertainty quantification.Paper: LUMA: A Benchmark Dataset for Learning from Uncertain and Multimodal Data
Code:… See the full description on the dataset page: https://huggingface.co/datasets/bezirganyan/LUMA.Danbooru-WD-EVA-EmbeddingsThis dataset includes WD EVA v2 large embeddings for danbooru images. Pixiv will be added later. Tensors under w are direct outputs and match indexes for WD EVA model. Tensors under e are from WD EVA as well however these strips the classifiaction head, they are smaller and suitable for deduplication computing for example.
The indexes of w, e and f (filename) match.
WD EVA Model: https://huggingface.co/SmilingWolf/wd-eva02-large-tagger-v3
lagenda_split
LAGENDA Dataset
This is a community mirror of the LAGENDA dataset created by LayerTeam. It has been uploaded here for easier access and integration with the Hugging Face datasets library.
All credit, rights, and accolades belong to the original authors. Please see the citation section below.
Dataset Description
LAGENDA (Large Age and Gender Dataset) is a dataset designed for age and gender recognition tasks. It addresses common biases in existing datasets by ensuring a… See the full description on the dataset page: https://huggingface.co/datasets/uaebn/lagenda_split.emnist-letters-tiny
Dataset Card for EMNIST-Letters-10k
A random subset of the train and test splits from the letters portion of EMNIST
This is a FiftyOne dataset with 10000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/emnist-letters-tiny.clearwrist-pediatric-wrist-xrayClearWrist: Pediatric Wrist Fracture X-Ray Dataset
20,327 labeled pediatric wrist radiographs, rebuilt from GRAZPEDWRI-DX with clean patient-level splits, verified fracture ground truth, and YOLO-style bounding box annotations.
Overview
This dataset packages the full GRAZPEDWRI-DX corpus, 20,327 pediatric wrist radiographs from 6,091 patients treated at the Department for Pediatric Surgery of the University Hospital Graz between… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/clearwrist-pediatric-wrist-xray.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.Lurcher_10x
Lurcher 10x Microscopy Dataset
Dataset overview
This dataset consists of 2-D microscopy images of histologically stained 3-D structures in tissue sections through the cerebellum of 21 mouse brains. Animals are grouped into wild-type controls (n = 10) and Lurcher mutant mice (n = 11). The classification task is to distinguish Lurcher mutant mice from wild-type controls.
All images were captured at low magnification (10x) and stained with Cresyl violet, a general… See the full description on the dataset page: https://huggingface.co/datasets/USF-CS-Microscopy-Image-Analysis/Lurcher_10x.LAION-Beyond
LAION-Beyond: Reproducible Vision-Language Models Meet Concepts Out of Pre-Training
📄 Paper (CVPR 2025) |
💻 Code |
🌐 Project Page
Dataset Summary
LAION-Beyond is the first multi-domain benchmark specifically designed to evaluate the Out-of-Pre-training (OOP) generalization of vision-language models (e.g., CLIP, OpenCLIP, EVA-CLIP).
We distinguish two types of visual concepts:
IP (In-Pre-training): concepts that appear in the pre-training data (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/MHuangX/LAION-Beyond.lfw
LFW HF-ready
This folder packages the local LFW (Labeled Faces in the Wild) images as a
Hugging Face imagefolder dataset with the canonical 10-fold verification
pairs file.
Layout
lfw/
├── README.md
├── pairs.csv
└── train/
├── images/<shard>/<file>.jpg
└── metadata.csv
metadata.csv columns
file_name: relative image path used by ImageFolder, e.g. images/000/Aaron_Eckhart_0001.jpg.
label: numeric identity label.
label_name / identity: identity name.… See the full description on the dataset page: https://huggingface.co/datasets/marcelohaps/lfw.MOT20
MOT20
MOT20 is a benchmark dataset for single-camera multi-object tracking (MOT) and pedestrian detection in very crowded real-world scenes. This Hugging Face repository provides MOT20 in the original MOTChallenge-style structure for research, benchmarking, training, and evaluation of multi-object tracking systems.
MOT20 was introduced to stress-test MOT methods in high-density pedestrian scenes, including crowded squares, indoor train stations, stadium exits, and pedestrian… See the full description on the dataset page: https://huggingface.co/datasets/Lekim89/MOT20.LADI-v2-dataset
Dataset Card for LADI-v2-dataset
Dataset Summary : v2
The LADI-v2 dataset is a set of aerial disaster images captured and labeled by the Civil Air Patrol (CAP). The images are geotagged (in their EXIF metadata). Each image has been labeled in triplicate by CAP volunteers trained in the FEMA damage assessment process for multi-label classification; where volunteers disagreed about the presence of a class, a majority vote was taken. The classes are:
bridges_any… See the full description on the dataset page: https://huggingface.co/datasets/MITLL/LADI-v2-dataset.GSD-Sensitivity-Taxonomy-Labels
GSD-Sensitivity Taxonomy: Task Labels for Remote Sensing VQA
Per-task D / M1 / M2 taxonomy labels, inter-annotator agreement (IAA) data, and
evaluation traces for four public RS-VQA benchmarks.
Companion to *G. Park and D.-H. Lee, "Identifying the Measurement Gap in Remote
Sensing VQA with a GSD-Sensitive Taxonomy," IEEE Geosci. Remote Sens. Lett., 2026*
— accepted, DOI to follow. Code: github.com/ganghyunnnn/GSD-Sensitivity-Taxonomy
⚠️ This dataset contains annotations and… See the full description on the dataset page: https://huggingface.co/datasets/ganghyunnnn/GSD-Sensitivity-Taxonomy-Labels.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.
