datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.TreeOfLife-10M
Dataset Card for TreeOfLife-10M
Dataset Summary
With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M.webvid-10Mimagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.imagenet-12k-wds
Dataset Summary
This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm.
The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k
NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.stable-diffusion-v1-5-glazed
Dataset Card for Stable Diffusion v1.5 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by runwayml/stable-diffusion-v1-5
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/stable-diffusion-v1-5-glazed.Lurcher_10x
Lurcher 10x Microscopy Dataset
Dataset overview
This dataset consists of 2-D microscopy images of histologically stained 3-D structures in tissue sections through the cerebellum of 21 mouse brains. Animals are grouped into wild-type controls (n = 10) and Lurcher mutant mice (n = 11). The classification task is to distinguish Lurcher mutant mice from wild-type controls.
All images were captured at low magnification (10x) and stained with Cresyl violet, a general… See the full description on the dataset page: https://huggingface.co/datasets/USF-CS-Microscopy-Image-Analysis/Lurcher_10x.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.casia-char-1
CASIA Character Sample Dataset
This dataset is adapted from CASIA Online and Offline Chinese Handwriting Databases,
but this only contains character level sample data (from the offline database). The first column is the ground truth label (single character from
GB2312 charset) and the second one is byte sequences of the decoded PNG files from the original .gnt files.
Conditions of Academic Use
Please refer to the official page for more information.
All samples in the… See the full description on the dataset page: https://huggingface.co/datasets/UndefinedCpp/casia-char-1.DomainNetData downloaded from WILDS (Download, paper, project).
This dataset contains some copyrighted material whose use has not been specifically authorized by the copyright owners. In an effort to advance scientific research, we make this material available for academic research. We believe this constitutes a fair use of any such copyrighted material as provided for in section 107 of the US Copyright Law. In accordance with Title 17 U.S.C. Section 107, the material on this site is distributed… See the full description on the dataset page: https://huggingface.co/datasets/wltjr1007/DomainNet.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.SAGE
SAGE: Scalable Agentic Grounded Evaluation for Crop Disease Diagnosis
Paper: arXiv:2605.09768
Arshad, Roy, Shen, Elango, Chiranjeevi, A. K. Singh, Ganapathysubramanian, Hegde, A. Singh, Sarkar.
~801,807 images across 335 crops / 1,251 disease classes, each enriched with
source-grounded visual symptom knowledge: every visual description is
extracted verbatim from an authoritative web page (university extension / APS /
CABI) with a source URL and a supporting quote, and… See the full description on the dataset page: https://huggingface.co/datasets/tirtho149/SAGE.cifar100-enrichedThe CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).EpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.planktonzilla-17M
Planktonzilla-17M Dataset
Overview
planktonzilla-17M is a large-scale, comprehensive dataset combining 17 million plankton images from all publicly available -to the best
of our knowledge- labeled plankton datasets. This unified collection enables researchers to train robust deep learning models for plankton
identification and classification across diverse imaging systems and oceanographic environments.
Each image includes a standardized taxonomic hierarchy… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/planktonzilla-17M.artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.OverMaps_1k
🗺️ OverMaps-1K Dataset
OverMaps-1K represents the convergence of classical 3D reconstruction rigor with the massive data requirements of modern Generative AI.Designed to bridge the gap between scale and quality, this dataset moves beyond the limitations of uncontrolled web videos or synthetic renders. It provides a foundation for training spatial reasoning in foundation models, advancing neural rendering, and developing robotic perception systems that can navigate… See the full description on the dataset page: https://huggingface.co/datasets/OverTheReality/OverMaps_1k.montgomery-shenzhen-tuberculosis-cxr
Montgomery and Shenzhen Tuberculosis Chest X-rays
This repository packages two public chest-radiograph collections for the accompanying educational notebooks on classification, detection, and segmentation. It preserves the original folder names and file layout so the notebooks can use the files directly.
Contents
ChinaSet_AllFiles/
├── CXR_png/
└── ClinicalReadings/
MontgomerySet/
├── CXR_png/
├── ClinicalReadings/
└── ManualMask/
├── leftMask/
└──… See the full description on the dataset page: https://huggingface.co/datasets/Famatsu123/montgomery-shenzhen-tuberculosis-cxr.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
Minecraft-Skins-Captioned-1M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical.
image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.nanopath-fairness-tiles
nanopath-fairness-tiles
Pre-tiled histopathology patches from CPTAC whole-slide images, used as the
external out-of-distribution validation set for a study on pretraining-time
vs. post-hoc fairness in histopathology foundation models.
Contents
Per-cohort folders, each slides_full/<slide_id>.parquet (one row per tile:
case_id, slide_id, tile_idx, image) + labels.tsv:
cohort
organ / task
slides
cptac_lung
NSCLC — LUAD vs LSCC subtype
604
cptac_gbm
GBM —… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/nanopath-fairness-tiles.Danbooru-2024-Filtered-1MGeometricShapeDataset
Geometric Shape Dataset
Introduction to Dataset
The Geometric Shape Dataset is a large-scale, synthetically generated computer vision dataset containing 1,680,000 instances of various geometric shapes and lines. It is designed for training image classification, feature extraction, and pattern recognition models across different scales.
The dataset is available in three distinct resolution configurations: 30x30, 40x40, and 50x50 (560,000 images per resolution).… See the full description on the dataset page: https://huggingface.co/datasets/OmerTurk1/GeometricShapeDataset.danbooru2024-latents-sdxl-1ktar
Danbooru 2024 SDXL VAE latents in 1k tar
Dedicated dataset to align deepghs/danbooru2024-webp-4Mpixel. "4MP-Focus" for average raw image resolution.
Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024.
Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py.
Used for kohya-ss/sd-scripts. In theory it may replace… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2024-latents-sdxl-1ktar.chest-xray-14-320
NIH Chest X-ray14 - 320x320 Processed for CheXVision
Project Resources
GitHub repository
Presentation deck
Live demo
Scratch model
DenseNet model
This dataset repackages the raw NIH Chest X-ray14 source dataset from
alkzar90/NIH-Chest-X-ray-dataset
into a data-only Parquet dataset for the CheXVision project.
Dataset Summary
Source format: 12 ZIP archives of original chest X-ray images plus CSV manifests
Output format: data-only Parquet shards under data/… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/chest-xray-14-320.showui-web-processed
ShowUI-Web Processed
Flattened, normalized, and scenario-split version of showlab/ShowUI-web.
Each row is a single (instruction, UI element) pair with normalized bounding-box coordinates.
Schema
Column
Type
Description
sample_id
string
Unique row identifier ({row}_{element})
screenshot_id
string
Groups elements from the same screenshot
image_relpath
string
Relative path to the screenshot image
scenario
string
Website/domain inferred from the image path… See the full description on the dataset page: https://huggingface.co/datasets/e1879/showui-web-processed.truck-images
Truck Detection and Counting Dataset
This repository contains multiple computer vision datasets for truck detection, counting, and classification.
1. Raw Truck Images (Root folder)
Number of images: 466
Format: Unannotated images (JPEG/PNG)
License: CC BY 4.0
2. Truck Counting Dataset
Number of images: 1769
Format: YOLOv8 format (images, labels, and data.yaml)
Location: truck_counting_yolov8/ folder
License: CC BY 4.0 (Roboflow export)
3. Trucks… See the full description on the dataset page: https://huggingface.co/datasets/shravya11/truck-images.pfm1-landmine-uav-vnir-hsi-IGARSS-2026
PFM-1 Landmine VNIR Hyperspectral Imaging Dataset (IGARSS 2026)
Dataset Description
This dataset 1 contains Visible and Near-Infrared (VNIR) Hyperspectral Imaging (HSI) data prepared for the
IEEE International Geoscience and Remote Sensing Symposium (IGARSS) 2026.
This specific version is a refined subset of the original benchmark dataset 2. While the original
release provided full radiance cubes, broad GCP/AeroPoint data, and reference ground… See the full description on the dataset page: https://huggingface.co/datasets/SagarLekhak/pfm1-landmine-uav-vnir-hsi-IGARSS-2026.color-multi-fractal-db-1k
Dataset Card for Color Multi Fractal DB 1k
This is a pre-generated 1k classes, 1M images colored-multi-fractal-images dataset based on Improving Fractal Pre-training by Connor Anderson et al. and Multi-Fractal-Dataset by FYSignate1009.
We have changed some fractal parameters so that our ViT pretraining can converge. Modified parameters can be found on this repo.
You can pretrain vision transformers without worrying about dataset licensing for commercial use.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/color-multi-fractal-db-1k.line-ex
LineEX
This dataset repo contains:
the uploaded train split with 397,993 images
the released test split with 20,000 images
Repo: 13point5/line-ex
Shared Schema
image_id
file_name
image
width
height
data_type
chart_elements
lines
chart_elements fields
annotation_id
category_id
category_name
bbox_xywh
area
text
line_id
lines fields
annotation_id
category_id
category_name
line_name
polyline_xy
raw_series_xy
area
Notes
The repo… See the full description on the dataset page: https://huggingface.co/datasets/13point5/line-ex.
