datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mnist
Dataset Card for MNIST
Dataset Summary
The MNIST dataset consists of 70,000 28x28 black-and-white images of handwritten digits extracted from two NIST databases. There are 60,000 images in the training dataset and 10,000 images in the validation dataset, one class per digit so a total of 10 classes, with 7,000 images (6,000 train images and 1,000 test images) per class.
Half of the image were drawn by Census Bureau employees and the other half by high school students… See the full description on the dataset page: https://huggingface.co/datasets/ylecun/mnist.Obshazard-bench
ObsCrisis-Bench
A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks.
Dataset Description
ObsCrisis-Bench contains 4,202 VQA samples across 127 extreme weather events in 8 disaster categories, covering 61 countries. Each sample combines satellite multispectral imagery (AMSU-A, HIRS, MHS sensors) with optional weather station data, and requires models to perform risk assessment, type classification, timing… See the full description on the dataset page: https://huggingface.co/datasets/YYQ898/Obshazard-bench.ARTO-Gen-Dataset
ARTO-KG: A Synthetic Artwork Dataset for Knowledge-Enhanced Understanding
Dataset Description
ARTO-KG is a large-scale synthetic artwork dataset that bridges visual content and structured knowledge through ontology-guided automated generation. Each artwork is annotated with comprehensive RDF knowledge graphs aligned with the ARTO ontology.
Dataset Summary
Total Artworks: 10,108 high-resolution images (1024×1024)
Object Instances: 39,878 (average… See the full description on the dataset page: https://huggingface.co/datasets/youngcan1/ARTO-Gen-Dataset.mixlora-eval-data
🚀 MixLoRA Evaluation Data
This dataset is the held-out multimodal evaluation suite used in
Multimodal Instruction Tuning with Conditional Mixture of LoRA (ACL 2024).
It bundles 9 instruction-formatted tasks (mm_tasks/) plus the MME benchmark
(mme/) used to evaluate MixLoRA and baseline models in the paper.
The 9 tasks in mm_tasks/ are the zero-shot / held-out task split from
Vision-Flan. MME is a
separate benchmark, evaluated independently.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/yingss/mixlora-eval-data.Glint360k
Dataset Card for Glint360K
Citiation by InsightFace Repository
We clean, merge, and release the largest and cleanest face recognition dataset Glint360K, which contains 17091657 images of 360232 individuals. By employing the Patial FC training strategy, baseline models trained on Glint360K can easily achieve state-of-the-art performance. Detailed evaluation results on the large-scale test set (e.g. IFRT, IJB-C and Megaface) are as follows:
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yayoimizuha/Glint360k.DamageTriage-Bench
DamageTriage-Bench
DamageTriage-Bench is a footprint-conditioned benchmark for per-building damage
typing from single post-event aerial images. Its five classes distinguish roof
from structural damage and partial from total affected extent:
ID
Class
0
Undamaged
1
Partial Roof Damage
2
Total Roof Damage
3
Partial Structural Damage
4
Total Structural Collapse
Quick statistics
Item
Value
Tiles
7,472 (1024 × 1024 PNG)
Labeled… See the full description on the dataset page: https://huggingface.co/datasets/Ymx1025/DamageTriage-Bench.Mirage-Test
🌊 Mirage-Test Dataset
Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models.
It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics.
The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity.
📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.glaucoma-expert-cot-final
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
glaucoma / not
train.jsonl
823
304 / 519
val.jsonl
92
46 / 46
test.jsonl
159
79 / 80
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train",
"final_diagnosis_GT":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-final.artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.MMS-VPR
MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments
Overview
MMS-VPR is the first large-scale multimodal street-level visual place recognition dataset featuring comprehensive integration of images, videos, and rich textual annotations with day–night coverage and a 7-year temporal span in dense pedestrian-only environments.
MMS-VPR comprises 110,529 images and 2,527 video clips… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.gelsight-mini-pretrain
GelSight Mini Pretrain
~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92.
Frames
Sources
Real
536K
FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad
Sim
317K
sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated)
NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.Urban-ImageNet
🏙️ Urban-ImageNet
A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception from Social Media Imagery.
Urban-ImageNet fills a critical gap between computer vision and urban studies by treating cities not simply as visual scenes, but as lived, socially produced, and experientially activated spaces.
Overview
ImageNet taught models to recognise objects. Urban-ImageNet teaches them to understand how people experience cities.… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/Urban-ImageNet.UnsafeBench
Dataset Card for Dataset Name
[Update]: we added the caption/prompt information (if there is one) in case other researchers need it. It is not used in our study though.
The dataset consists of 10K safe/unsafe images of 11 different types of unsafe content and two sources (real-world VS AI-generated).
Dataset Details
Source
# Safe
# Unsafe
# All
LAION-5B (real-world)
3,228
1,832
5,060
Lexica (AI-generated)
2,870
2,216
5,086
All
6,098
4,048
10,146… See the full description on the dataset page: https://huggingface.co/datasets/yiting/UnsafeBench.HCSU
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding (V1.0)
ECCV 2026
Yinsheng Yao*, Yan Liu*, Chen Ye†
* Equal contribution. † Corresponding author.
Abstract: Addressing the "Knowledgeable but unperceptive" dilemma where existing Large Vision-Language Models (LVLMs) possess historical knowledge but lack fine-grained calligraphy style perception, we introduce HCSU—the first large-scale dataset and evaluation benchmark specifically… See the full description on the dataset page: https://huggingface.co/datasets/YAN-LIU05/HCSU.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/yh123yh/NCT-CRC-HE.glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/Varsha-Y12/NCT-CRC-HE.Nurisk
Nurisk: VQA for Risk Assessment in Autonomous Driving
Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains:
image: a BEV image
question: a driving-related question
answer: the ground truth answer
Paper
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 .
Framework
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Yuan-avs/Nurisk.exp03-l23-hardening
Exp 03 / 03b — Klüver L2/3 hardening study (SDXL + SD 3.5)
Pre-registered study from the Operating System Hypothesis project.
Sweep classifier-free guidance across two architectures and score every output
blind, on two independent rubrics, for how far object structure has come apart.
The prediction was written down and committed before the run. The commit
dates in the GitHub repo are the proof.
What is here that is not on GitHub
The 860 generated PNGs. Every text… See the full description on the dataset page: https://huggingface.co/datasets/youssefhassan13/exp03-l23-hardening.GAMMA
GAMMA — Glaucoma grading from Multi-Modality imAges (Challenge dataset)
Image: Dataset Samples.
Short description
GAMMA is the first public multi-modality glaucoma grading dataset that pairs 2D color fundus photographs with 3D OCT volumes for each sample. It was released as part of the GAMMA challenge (OMIA8 / MICCAI 2021) to encourage algorithms that combine fundus and OCT information for… See the full description on the dataset page: https://huggingface.co/datasets/yujiaxue/GAMMA.apextrack-track-condition-dataset
ApexTrack Balanced Racing Track Condition Dataset
This dataset contains racing track condition images curated and balanced across three primary condition classes: DRY, DAMP, and WET.
It serves as the official training, validation, and evaluation benchmark for the ApexTrack AI Vision Transformer model (yuvrajengines/apextrack-track-condition-v2).
Dataset Structure
apextrack_balanced/
├── train/
│ ├── damp/ (100 images)
│ ├── dry/ (100 images)
│ └── wet/… See the full description on the dataset page: https://huggingface.co/datasets/yuvrajengines/apextrack-track-condition-dataset.CHUBS
CHUBS: A Large-Scale Dataset of Chu Bamboo Slip Script
Code | Paper (upcoming)
Introduction
This is a large-scale dataset of Chu bamboo slip (CBS, Chinese: 楚简, chujian) script, an ancient Chinese script used during the Spring and Autumn period over 2,000 years ago. This dataset consists of two parts:
The main dataset where each example is an image and the corresponding text label. This part is contained in the glyphs.zip ZIP file.
A character detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/chen-yingfa/CHUBS.CIFAKE_autotrain_compatible
Dataset Card for CIFAKE_autotrain_compatible
Dataset Summary
This is a copy of the CIFAKE dataset created by Dr Jordan J. Bird and Professor Ahmad Lotfi. See more information on the original data card on Kaggle.
The real images used are from CIFAR-10. The fake images were created by the authors using Stable Diffusion v1.4.
This dataset removes the train/test structures in the original dataset to allow compatibility with HuggingFace's AutoTrain. It removes the test split… See the full description on the dataset page: https://huggingface.co/datasets/yanbax/CIFAKE_autotrain_compatible.AniPro
AniPrO: Interpretable Anime Image Provenance Detection via Multi-Dimensional Semantic Reasoning
CGI 2026
Yan Liu, Baoxiang Huang, Zi'an Wang, Wenbo Xie
Tongji University
AniPrO is a balanced diagnostic benchmark for interpretable anime image provenance detection. It distinguishes human-drawn illustrations, AI-inpainted images, and text-to-image generations through observable visual evidence rather than source labels alone.
This repository contains the released image… See the full description on the dataset page: https://huggingface.co/datasets/YAN-LIU05/AniPro.low-guidance-cfg-sweep
Sub-CFG guidance sweep (g = 0 → 2), SDXL + SD 3.5
Exploratory. Not pre-registered. Not a result.
No hypothesis was committed before these runs, there is no pre-specified
statistical model, and no p-values are reported anywhere in this dataset.
The sibling Exp 03 dataset
is pre-registered, with commit dates as proof. This one is not. Treat it
as a reason to design an experiment, not as evidence for a claim.
From the Operating System Hypothesis project. Exp 01 and Exp 03 both… See the full description on the dataset page: https://huggingface.co/datasets/youssefhassan13/low-guidance-cfg-sweep.toy-car-annotation-YOLOHey everyone,
In my final year project, I created Smart Traffic Management System.The project was to manage traffic lights' delays based on the number of vehicles on road.I made everything worked using Raspberry Pi and pre-recorded videos but it was a "final year project", it was needed to be tested by changing videos frequently which was a kind of hustle. Collecting tons of videos and loading them in Pi was not too hard but it would have cost time, by every time changing names of videos in… See the full description on the dataset page: https://huggingface.co/datasets/tubasid/toy-car-annotation-YOLO.TurkishFoods-25
⚠️ Important: English version is available below.
TürkSofrası-25 (TurkishFoods-25) Veri Seti
TürkSofrası-25, 25 farklı geleneksel Türk yemeğine ait toplam 11.461 görsel içeren ve yemek tanıma/sınıflandırma amaçlı hazırlanmış bir görüntü veri setidir. Görseller .jpg formatında olup her sınıf için ayrı klasörlerde yer almaktadır.
Veri seti, Hugging Face datasets kütüphanesi biçimindedir ve image (görsel) ile label (etiket) olmak üzere iki özelliğe sahiptir. Etiketler class_label… See the full description on the dataset page: https://huggingface.co/datasets/yunusserhat/TurkishFoods-25.ARGUS_DATASET
🌍 ARGUS DATASET
Multi-Domain Global Landmark, Streetscape & Geospatial Intelligence Dataset
📌 Dataset Overview
ARGUS_DATASET is an open, research-grade geospatial intelligence (GEOINT), computer vision, and visual geolocation benchmark dataset. It provides verified, multi-angle landmark photography, panoramic street-level imagery, spatial index databases, and DCT perceptual hash trees across sovereign nations, territories, and municipalities… See the full description on the dataset page: https://huggingface.co/datasets/YTxFSGAMERz/ARGUS_DATASET.
