datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.OpenFake
Dataset Card for OpenFake
Known issues
Prompt–image misalignment in the synthetic split (reported November 2025, fix pending)
For five of the eighty generators, the prompt field attached to synthetic
images does not correspond to the prompt actually used to generate that image.
Affected generators:
flux-realism
sd-3.5
sdxl-realvis-v5
sd-1.5-dreamshaper
sd-1.5-epicdream
This affects approximately 19.77% of synthetic images. It was first reported in
discussion… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.TreeOfLife-200M
Dataset Card for TreeOfLife-200M
If you are looking for the original release TreeOfLife-200M dataset, as used in training BioCLIP 2 and presented the paper, please see Revision a8f38b4. The dataset, as presented here, was used to train BioCLIP 2.5 Huge; it completes the dataset cleaning process and resolves an issue where Observation.org occurrences were not included in the training data.
With 233 million images representing 933,798 taxa across the tree of life, TreeOfLife-200M… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M.fish-vista
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista)
Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images.
See Example Code to Use the Segmentation Dataset
Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset.
Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
oxford-iiit-pet
The Oxford-IIIT Pet Dataset
Description
A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting.
This instance of the dataset uses standard label ordering and includes the standard train/test splits. Trimaps and bbox are not included, but there is an image_id field that can be used to reference those annotations from official metadata.
Website: https://www.robots.ox.ac.uk/~vgg/data/pets/… See the full description on the dataset page: https://huggingface.co/datasets/timm/oxford-iiit-pet.femnist
Dataset Card for FEMNIST
The FEMNIST dataset is a part of the LEAF benchmark.
It represents image classification of handwritten digits, lower and uppercase letters, giving 62 unique labels.
Dataset Details
Dataset Description
Each sample is comprised of a (28x28) grayscale image, writer_id, hsf_id, and character.
Curated by: LEAF
License: BSD 2-Clause License
Dataset Sources
The FEMNIST is a preprocessed (in a way that resembles preprocessing for… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/femnist.GroMo25
GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting
Dataset Summary
GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.MIC21
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18029/
MIC21
Original description
One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is constantly expanding, still there is a huge demand for more datasets in terms of variety of domains and object classes covered. The goal of the… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/MIC21.webvid-10Msvg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.cola
COLA: Compose Objects Localized with Attributes
Self-contained Hugging Face port of the COLA benchmark from the paper
"How to adapt vision-language models to Compose Objects Localized with Attributes?".
📄 Paper: https://arxiv.org/abs/2305.03689
🌐 Project page: https://cs-people.bu.edu/array/research/cola/
💻 Original code & data: https://github.com/ArijitRay1993/COLA
This repository bundles the benchmark annotations as Parquet files and the referenced
images as regular files… See the full description on the dataset page: https://huggingface.co/datasets/array/cola.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.svgrepo
Dataset Card for SVGRepo Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGRepo.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under a specific open-source or permissive license, clearly indicated in its metadata. The SVG… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/svgrepo.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.REPID
REPID: Rendering Evaluation of Photographic Image Dataset
REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality.
Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.BioTrove
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove-Train dataset card on HuggingFace to access the samller BioTrove-Train dataset (40M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove.britannica-illustrated-pages
Britannica Illustrated Pages
115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition
(1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes
(838 Internet Archive items). A second config carries the classifier
score, OCR word count and provenance for every one of the 975,345 pages.
Two things the scan showed:
82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and
engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.amazon-berkeley-objects
Amazon Berkeley Objects (ABO)
A Hugging Face packaging of the Amazon Berkeley Objects (ABO) dataset. The
data content is the official CC BY 4.0 release from
https://amazon-berkeley-objects.s3.amazonaws.com/index.html. This mirror
changes only the packaging: files are grouped into typed Parquet shards, and
every original media file is preserved byte-for-byte and never transcoded.
Images use the datasets Image() feature, 3D product models use the native
Mesh() feature (original… See the full description on the dataset page: https://huggingface.co/datasets/suvadityamuk/amazon-berkeley-objects.imagenet-22k-wds
Dataset Summary
This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds)
This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.SiliciclasticReservoirs
Siliciclastic Reservoirs
Released by SciLM.ai: https://www.scilm.ai
1,000,000 synthetic 3D siliciclastic-reservoir geology cubes generated from rule-based sedimentological simulations (turbidite lobes + 6 fluvial-channel architectures + delta-fan distributary). Cubes are voxelized at (64, 64, 32) cells. Each sample carries facies, porosity, permeability, and a structured set of geological conditioning parameters.
Designed to train conditional generative models (flow matching… See the full description on the dataset page: https://huggingface.co/datasets/SciLM/SiliciclasticReservoirs.glint360k-wds-gz
Glint360K
This dataset is introduced in the Partial FC paper https://arxiv.org/abs/2010.05222.
There are 17,091,657 images and 360,232 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_. The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 1,385 shards in total.… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/glint360k-wds-gz.Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.invasive_plants_hawaii
Dataset Card for Invasive Plants Project
This dataset is aimed at the image multi-classification and segmentation of various leaf damage types caused by biocontrol agents. The dataset contains images of both the dorsal and ventral side of Clidemia Hirta leaves, that were all collected in January 2025 near Hilo (Hawaii), in dirt trails along Steinback Highway. Clidemia Hirta is a highly invasive plant on the island of Hawaii (Big Island).
Dataset Configurations and… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/invasive_plants_hawaii.PlantVillage
PlantVillage Dataset
The PlantVillage Dataset is an open access repository of 54,306 images of healthy and diseased plant leaves, collected to advance research in automated plant disease diagnosis. It covers 14 crop species and 26 diseases.
This dataset was introduced in the paper "Using Deep Learning for Image-Based Plant Disease Detection" by Mohanty et al. (2016).
Quick Start
The dataset comes with pre-defined 80/20 train/test splits that preserve the leaf grouping… See the full description on the dataset page: https://huggingface.co/datasets/mohanty/PlantVillage.eurosatRedistributed without modification from https://github.com/phelber/EuroSAT.
EuroSAT100 is a subset of EuroSATallBands containing only 100 images. It is intended for tutorials and demonstrations, not for benchmarking.
imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.
