datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/LiamLian0727/UIIS.laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.lfw
LFW HF-ready
This folder packages the local LFW (Labeled Faces in the Wild) images as a
Hugging Face imagefolder dataset with the canonical 10-fold verification
pairs file.
Layout
lfw/
├── README.md
├── pairs.csv
└── train/
├── images/<shard>/<file>.jpg
└── metadata.csv
metadata.csv columns
file_name: relative image path used by ImageFolder, e.g. images/000/Aaron_Eckhart_0001.jpg.
label: numeric identity label.
label_name / identity: identity name.… See the full description on the dataset page: https://huggingface.co/datasets/marcelohaps/lfw.efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.Lurcher_10x
Lurcher 10x Microscopy Dataset
Dataset overview
This dataset consists of 2-D microscopy images of histologically stained 3-D structures in tissue sections through the cerebellum of 21 mouse brains. Animals are grouped into wild-type controls (n = 10) and Lurcher mutant mice (n = 11). The classification task is to distinguish Lurcher mutant mice from wild-type controls.
All images were captured at low magnification (10x) and stained with Cresyl violet, a general… See the full description on the dataset page: https://huggingface.co/datasets/USF-CS-Microscopy-Image-Analysis/Lurcher_10x.clearwrist-pediatric-wrist-xrayClearWrist: Pediatric Wrist Fracture X-Ray Dataset
20,327 labeled pediatric wrist radiographs, rebuilt from GRAZPEDWRI-DX with clean patient-level splits, verified fracture ground truth, and YOLO-style bounding box annotations.
Overview
This dataset packages the full GRAZPEDWRI-DX corpus, 20,327 pediatric wrist radiographs from 6,091 patients treated at the Department for Pediatric Surgery of the University Hospital Graz between… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/clearwrist-pediatric-wrist-xray.GSD-Sensitivity-Taxonomy-Labels
GSD-Sensitivity Taxonomy: Task Labels for Remote Sensing VQA
Per-task D / M1 / M2 taxonomy labels, inter-annotator agreement (IAA) data, and
evaluation traces for four public RS-VQA benchmarks.
Companion to *G. Park and D.-H. Lee, "Identifying the Measurement Gap in Remote
Sensing VQA with a GSD-Sensitive Taxonomy," IEEE Geosci. Remote Sens. Lett., 2026*
— accepted, DOI to follow. Code: github.com/ganghyunnnn/GSD-Sensitivity-Taxonomy
⚠️ This dataset contains annotations and… See the full description on the dataset page: https://huggingface.co/datasets/ganghyunnnn/GSD-Sensitivity-Taxonomy-Labels.watercolour-rollouts-judge-led
Watercolour rollouts, judge-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 861 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the run with the original reward mix from the write-up, where the pairwise judge and
its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.AI-GenBench-fake_part
AI-GenBench: A New Ongoing Benchmark for AI-Generated Image Detection
Important: this is the fake part of the AI-GenBench dataset. To re-create the original benchmark, which includes real images, please check the official repository.
Important 2: before using, please check the licensing terms of the images included!
Details
The rapid advancement of generative AI has revolutionized image creation, enabling high-quality synthesis from text prompts while raising critical… See the full description on the dataset page: https://huggingface.co/datasets/lrzpellegrini/AI-GenBench-fake_part.liquidrandom-data
liquidrandom-data
Diverse seed data for ML/LLM training data generation pipelines.
Used by the liquidrandom Python package.
Dataset Summary
This dataset contains 520,080 seed data samples across 24 categories,
generated using a hierarchical taxonomy tree approach with LLM-based quality validation
and fuzzy deduplication. Data is stored as Parquet with zstd compression.
Categories
Category
Samples
File
Coding Tasks
30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.Spatial-DISE
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
📋 Overview
Spatial-DISE is a comprehensive benchmark dataset designed to evaluate spatial reasoning capabilities in vision-language models. The dataset focuses on various aspects of spatial intelligence including 3D perception, spatial transformation, and geometric reasoning across multiple difficulty levels.
🧪 Evaluation Support
Supported:… See the full description on the dataset page: https://huggingface.co/datasets/TACPS-liv/Spatial-DISE.Latent-Earth
Latent Earth: An Atlas of Architecture in Flux.2
200,000 images of 40,000 places on Earth, each rendered by a single
image model in a single state of its training, with five internal
representations recorded for every image while it was being generated.
Nothing else enters. Each prompt contains only a place's name; no
photographs, no maps, no climate records correct what the model proposes.
This is therefore not a depiction of the world but a probe of the model: a
survey of what… See the full description on the dataset page: https://huggingface.co/datasets/Punktiert/Latent-Earth.BigEarthNetV2-LMDB
TU Berlin
RSiM
DIMA
BigEarth
BIFOLD
reBEN (pre-converted to LMDB)
⚠️ Unofficial mirror. This is an unofficial, community-providedpre-conversion of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take precedence. Please refer to the authoritative… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-LMDB.llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.bakkhali-river-high-low-tide
Bakkhali River — High Tide vs Low Tide, Bangladesh
517 photographs of the Bakkhali River near Cox's Bazar, Bangladesh, documenting the same general stretch of river at high tide (264 images) and low tide (253 images). Captured across 10 separate sessions between 2 July and 15 August 2026.
This is not a frame-by-frame matched pair set — sessions were shot on different dates and the camera position varies within each session — but high- and low-tide frames come from the same short… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/bakkhali-river-high-low-tide.localizer
Dartbrains Localizer Dataset
A subset of the Brainomics/Localizer functional MRI dataset, prepared for the Dartbrains neuroimaging course at Dartmouth College.
Quick Start
Load beta maps (recommended for most exercises)
from datasets import load_dataset
ds = load_dataset("dartbrains/localizer", "betas")
img = ds[0]["nifti"] # nibabel.Nifti1Image
subject = ds[0]["subject"] # "S01"
condition = ds[0]["condition"] # "audio_computation"… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/localizer.oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.AEGISThis repository contains the data of the paper [AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images]
lamoda-fashion-product-images
High-Resolution Fashion Product Images
This dataset is a highly optimized, high-resolution subset of the popular Fashion Product Images Dataset originally hosted on Kaggle.
It contains thousands of unique e-commerce fashion products, combining high-resolution product images with multiple descriptive label attributes.
All low-resolution thumbnails and anomalies have been aggressively filtered out. Every image in this dataset has a minimum resolution of 640px on its shortest… See the full description on the dataset page: https://huggingface.co/datasets/PestoRosso/lamoda-fashion-product-images.meowcat-predictions
MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC
Per-pixel cell-type predictions generated by MeowCat on
H&E whole-slide images from two public cohorts:
Cohort
Tissue
Samples
h5ad payload
TCGA-LUAD
Lung adenocarcinoma
531
~60 GB
CPTAC-CCRCC
Clear-cell renal cell carcinoma
831
~93 GB
File layout
composition.parquet # long format: sample × cell_type → count, fraction
metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.watercolour-rollouts-hps-led
Watercolour rollouts, hps-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 872 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the middle point of the project's three reward mixes: the generic preference model
holds most of the weight, the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-led.the-un-laion-templeAll files uploaded. Enjoy!
Dataset Card for The Unlaion Temple
Dataset Details
Dataset Description
Laion-5B is still not public, so we decided to create our own dataset.
The Unlaion Temple is a raw dataset of CommonCrawl images (Estimated to be a total of 2 Billion urls). We haven't verified whether the links in this dataset are functional.
You are responsible for handling the data.
We've made some improvements to the dataset based on user feedback:
All… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Chiharu/the-un-laion-temple.
