datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fish-vista
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista)
Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images.
See Example Code to Use the Segmentation Dataset
Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset.
Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.GroMo25
GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting
Dataset Summary
GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.REPID
REPID: Rendering Evaluation of Photographic Image Dataset
REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality.
Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.britannica-illustrated-pages
Britannica Illustrated Pages
115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition
(1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes
(838 Internet Archive items). A second config carries the classifier
score, OCR word count and provenance for every one of the 975,345 pages.
Two things the scan showed:
82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and
engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.amazon-berkeley-objects
Amazon Berkeley Objects (ABO)
A Hugging Face packaging of the Amazon Berkeley Objects (ABO) dataset. The
data content is the official CC BY 4.0 release from
https://amazon-berkeley-objects.s3.amazonaws.com/index.html. This mirror
changes only the packaging: files are grouped into typed Parquet shards, and
every original media file is preserved byte-for-byte and never transcoded.
Images use the datasets Image() feature, 3D product models use the native
Mesh() feature (original… See the full description on the dataset page: https://huggingface.co/datasets/suvadityamuk/amazon-berkeley-objects.FakeCOCO
FakeCOCO dataset
Using 10 SOTA text-to-image models to generate fake images based on COCO captions
over 1M images
These models include:
SD15
SD21
SDXL
SD3
Playground2.5
PixArt alpha
PixArt sigma
unidiffuser
Flux.1
Stable Cascade
zju-eye-pretrain
ZJU Eye-Pretrain Dataset
Unified multi-source ophthalmological imaging dataset for foundation model pretraining and downstream tasks.
1.18M images spanning 57 cohorts with a strict 41-column unified manifest schema.
Composition
1,179,458 images · 57 cohorts · 8 modalities, under one strict 41-column unified manifest schema. Stored as HF parquet shards in data/{private_topcon, public_fundus, public_oct, public_new}/; loadable by batch or by single cohort (57… See the full description on the dataset page: https://huggingface.co/datasets/MaybeRichard/zju-eye-pretrain.zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.PP
SteelBench: A Diagnostic Benchmark for Vision-Language Models in Industrial Safety Monitoring
SteelBench is a diagnostic benchmark of densely annotated CCTV clips from an
operating integrated steel plant. It is designed to evaluate vision-language
models (VLMs) on real-world industrial action recognition, PPE assessment,
and safety-violation detection — under naturally occurring degradation
(dust, glare, steam, low light), at distances and crowdedness levels that
curated… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingHub/PP.openbrush
OpenBrush-75K
A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research.
Dataset Description
OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush.foodv2
Turkish Food & Nutrition Vision Dataset (Food v2)
High-resolution, human-POV Turkish food and packaged snack dataset with canonical nutrition macros and portion sizes.
📊 Dataset Summary
Repository: ozertuu/foodv2
Canonical Foods: 29,840
Images Total: 89,520
Average Images / Food: ~3.0
Image Sources: Yemeksepeti / DeliveryHero, GetirYemek, Migros Sanalmarket, Getir, Nefis Yemek Tarifleri, OpenFoodFacts.
🍽️ Features & Schema
image: PIL Image /… See the full description on the dataset page: https://huggingface.co/datasets/ozertuu/foodv2.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.SwissCrop25
SwissCrop25
A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2
time series, daily temperature data, and parcel-level crop type labels across seven growing
seasons (2019–2025).
Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
(TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team]
Highlights
Nationwide coverage of Switzerland (41,285 km²)
Seven growing seasons (2019–2025)
73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.danbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.GeoMeld
🌍 GeoMeld Multi-Modal Earth Observation Dataset (WebDataset)
GeoMeld is a large-scale multi-modal remote sensing dataset introduced in our CVPRW 2026 paper on semantically grounded foundation modeling.
GeoMeld contains approximately 2.5 million spatially aligned samples spanning heterogeneous sensing modalities and spatial resolutions, paired with semantically grounded captions generated through an agentic pipeline.
The dataset is designed to support multimodal representation… See the full description on the dataset page: https://huggingface.co/datasets/vimageiitb/GeoMeld.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.GTSRB
Dataset Card for GTSRB
Dataset Summary
The German Traffic Sign Benchmark is a multi-class, single-image classification challenge held at the International Joint Conference on Neural Networks (IJCNN) 2011. We cordially invite researchers from relevant fields to participate: The competition is designed to allow for participation without special domain knowledge. Our benchmark has the following properties:
Single-image, multi-class classification problem
More than 40… See the full description on the dataset page: https://huggingface.co/datasets/bazyl/GTSRB.relaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.TASTE
TASTE: Human Preferences for Design-Quality Image Comparison
This dataset is the human-evaluation corpus released alongside the
TASTE preference model. It contains panel rankings of generated
images across multiple quality dimensions — both aesthetic (does the
image look good?) and description-faithfulness (does the image match
what the prompt describes?) — plus a per-image hallucination
judgement.
Quick stats
Table
Rows
Notes
prompts.parquet
~200
one… See the full description on the dataset page: https://huggingface.co/datasets/purvanshi/TASTE.Icarus-dataset
Icarus
A unified multi-modal curriculum dataset for evolutionary neural architecture search. Every row is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. Support is non-empty for every task. Encoders read the Field descriptor (axes, value_type, n_classes, value_range, mask); mask is True where a value is padding/ignored. meta.class_names, when… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/Icarus-dataset.EpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.PovertyMap
PovertyMap-Wilds: Poverty mapping across different countries
Description
This is a processed version of LandSat 5/7/8 satellite imagery originally from Google Earth Engine under the names LANDSAT/LC08/C01/T1_SR,LANDSAT/LE07/C01/T1_SR,LANDSAT/LT05/C01/T1_SR,
nighttime light imagery from the DMSP and VIIRS satellites (Google Earth Engine names NOAA/DMSP-OLS/CALIBRATED_LIGHTS_V4 and NOAA/VIIRS/DNB/MONTHLY_V1/VCMSLCFG)
and processed DHS survey metadata obtained from… See the full description on the dataset page: https://huggingface.co/datasets/1aurent/PovertyMap.questFish2024
Dataset Card for QUEST Fish 2024
Images collected by teachers during a QUEST workshop. In 2024, the images were of fish collected from bodies of water near Princeton University.
Dataset Details
Dataset Structure
/dataset/
<folder>/
<img_id 1>.png
<img_id 2>.png
...
<img_id n>.png
...
<img_id 1>.png
<img_id 2>.png
...
<img_id n>.png
fieldData2024.csv
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/questFish2024.liquidrandom-data
liquidrandom-data
Diverse seed data for ML/LLM training data generation pipelines.
Used by the liquidrandom Python package.
Dataset Summary
This dataset contains 520,080 seed data samples across 24 categories,
generated using a hierarchical taxonomy tree approach with LLM-based quality validation
and fuzzy deduplication. Data is stored as Parquet with zstd compression.
Categories
Category
Samples
File
Coding Tasks
30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.
