datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.amazon-berkeley-objects
Amazon Berkeley Objects (ABO)
A Hugging Face packaging of the Amazon Berkeley Objects (ABO) dataset. The
data content is the official CC BY 4.0 release from
https://amazon-berkeley-objects.s3.amazonaws.com/index.html. This mirror
changes only the packaging: files are grouped into typed Parquet shards, and
every original media file is preserved byte-for-byte and never transcoded.
Images use the datasets Image() feature, 3D product models use the native
Mesh() feature (original… See the full description on the dataset page: https://huggingface.co/datasets/suvadityamuk/amazon-berkeley-objects.zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.SwissCrop25
SwissCrop25
A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2
time series, daily temperature data, and parcel-level crop type labels across seven growing
seasons (2019–2025).
Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
(TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team]
Highlights
Nationwide coverage of Switzerland (41,285 km²)
Seven growing seasons (2019–2025)
73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.self-driving-GTA-V
Self Driving GTA V Dataset
Dataset Varients
Mini : Link
Training Data(1-100) : Link
Training Data(101-200) : Link
Info
Image Resolution : 270, 480
Mode : RGB
Dimension : (270, 480, 3)
File Count : 100
Size : 1.81 GB/file
Total Data Size : 362 GB
Total Frames : 1 Million
Data Set sizes
Mini :
Folder Name : mini
Files : 01
Total Size : 1.81 GB
Total Frames : 5000
First Half
Folder Name : Training Data(1-100)
Files :… See the full description on the dataset page: https://huggingface.co/datasets/sartajbhuvaji/self-driving-GTA-V.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
Accepted to CVPRW 2026.
GitHub | CVF Open Access | arXiv | Model
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.stem-diagrams
STEM Diagrams
30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures)
extracted from arXiv papers across six engineering fields, each with a source
attribution and a quality score. Built by an LLM-curated pipeline and used to show
that a small frozen-feature classifier can replace the paid LLM labeling gate.
Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026)
Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.similar-but-different
Similar But Different — Sentinel-2 patches with deceptive RGB
30,927 32×32 Sentinel-2 L2A multispectral patches (12 bands resampled to
10 m) across ten ESA WorldCover classes, selected so that the visible
bands are uninformative by construction: every patch sits in a region of
RGB-mean space dominated by patches of other classes, while its
NIR / red-edge / SWIR response stays class-informative.
The dataset is a controlled probe for one question: does a model actually
use the… See the full description on the dataset page: https://huggingface.co/datasets/calebrob6/similar-but-different.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.drawvla-prompt-validation-clean
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-clean.dental-implant-surgery-sample
Dental Implant Surgery — Multimodal Annotated Video (Sample Case)
A public sample from one complete All-on-4 full-arch mandibular dental implant
surgery: two synchronised camera angles, the operating surgeon narrating while
he works, and six layers of structured clinical annotation (L0–L5) tied frame by
frame to what he said.
This is a showcase slice, not the whole case. What is here is enough to judge
the structure, the annotation quality and the honesty of the documentation.… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.WildFake-Sample
WildFake-Sample
A 30,000-image sample of WildFake (Hao et al., AAAI 2025,
arXiv:2402.11843;
original dataset),
covering generators and real-image sources outside DDA/SID — a held-out
generalization slice, not a copy of the full ~3.6M-image dataset. All credit
for the images goes to WildFake's original authors. Built for
Buxt-Codes/AIGI-Detection
(branch LoRC-PC) — see that repo's HANDOFF.md for the evaluation
methodology and results.
Composition
Fake (19,500):… See the full description on the dataset page: https://huggingface.co/datasets/buxtcodes/WildFake-Sample.drawvla-prompt-validation
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation.drawvla-prompt-validation-v3
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-v3.msc-imagenet100
MSC · ImageNet-100
Per-sample Minimum Sufficient Compute for ImageNet-100 at 224px: the
cost-normalised compute each sample needs before its decision has
settled. Companion to the CIFAR-100 study at
Shanmuk4622/msc-cifar100.
Seed reliability (rho_seed, tau=0.1)
architecture
family
seeds
rho_seed
Jaccard@10
top-1
resnet50
resnet
2
0.8220
0.6208
0.8237
vit_small_p16
vit
2
0.6492
0.2777
0.6079
rho_seed is the Spearman correlation between the MSC of… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/msc-imagenet100.image-aesthetic-scores
Rule34.nexus · Licence: Rule34.nexus Derived Dataset Licence 1.0
Rule34.nexus Image Aesthetic Scores
1. Overview
This dataset contains per-image aesthetic predictions for images in the Rule34.nexus corpus.
Predictions were generated using
discus0434/aesthetic-predictor-v2-5. Source images are not
included in this dataset — only opaque post identifiers, the source image's SHA-256 hash,
the post's content type, and the predicted score.… See the full description on the dataset page: https://huggingface.co/datasets/rule34nexus/image-aesthetic-scores.africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all
Brain Tumor (MRI) Detection Colourized with EHR | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all.MONITRS
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
Dataset Description
Paper: NeurIPS 2025 (Spotlight)
Contact: revankar@cs.cornell.edu
MONITRS contains ~10,000 FEMA disaster events with temporal Sentinel-2 satellite imagery, natural language captions from news articles, geotagged locations, and question-answer pairs for disaster monitoring research.
Supported Tasks
Event classification
Temporal grounding
Location grounding
Visual… See the full description on the dataset page: https://huggingface.co/datasets/ShreelekhaR/MONITRS.rvl-cdip-filtered
RVL-CDIP Filtered Dataset
This dataset contains filtered images from the RVL-CDIP dataset, focusing on 4 specific document types.
Dataset Summary
A filtered subset of the RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset containing 100,000 images across 4 document categories. Each image is stored as base64-encoded data in Parquet format for efficient processing.
Classes
Label
Class Name
Description
0
letter
Personal and… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/rvl-cdip-filtered.sd2-staged-foreign-objects
SD2 staged-laboratory foreign-object frames (heydonto)
97 frames · 163 frame-level annotation rows · 8 staged laboratory takes · CC BY 4.0
This dataset discloses and carries the SD2 own-footage portion of the training data of the ORena SAVE FOCUS challenge entry's FRAME-track component: 163 of that component's 58,086 pooled training rows. The other sources of that corpus are not part of this dataset.
What is in it
frames/ — 97 JPEG frames (filename = first 16 hex… See the full description on the dataset page: https://huggingface.co/datasets/HeyDonto/sd2-staged-foreign-objects.fish-vista
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista)
Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images.
See Example Code to Use the Segmentation Dataset
Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset.
Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/saffatgazi/fish-vista.Mr.Porter.Product.prices.Sweden
Mr Porter web scraped data
About the website
The EMEA region, specifically Sweden, has seen a significant rise in the luxury online retail industry, where Mr Porter operates. The growth has primarily been driven by the fast-paced digitalization, significant internet penetration, and a growing number of digitally native consumers. Additionally, Swedish consumers, renowned for their fashion-forward approach, have demonstrated a strong appetite for luxury fashion products… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Sweden.Prada.Product.prices.Sweden
Prada web scraped data
About the website
The Luxury Fashion Industry in the EMEA region, particularly in Sweden, is a thriving market with high demand for exclusive and high-end products. Prada, a renowned player in this industry, holds a significant presence. The industry is currently experiencing a significant shift towards digitalization and online retail, also known as Ecommerce, fueled by changing consumer behaviors and advancements in technology. A concrete example… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Prada.Product.prices.Sweden.kamari-safe-open-v0
Kámárí-Safe Open v0 (benchmark)
A frozen, leakage-free benchmark for African-tailored age verification. It holds manifests and
split tables, not raw images (paths, hashes, labels, skin band, quality). Use it to measure age
accuracy and, more importantly, child-safety.
Headline metric
Minor-Pass-Through Rate (MPTR) is the headline: the fraction of true minors a model passes as
adults, reported overall, at 21, and for dark + brown skin. Report MPTR alongside MAE; a… See the full description on the dataset page: https://huggingface.co/datasets/Shinzmann/kamari-safe-open-v0.StatMapCorpus
StatMapCorpus v1
23,549 English-language statistical maps identified in MapPool, annotated for cartographic
method by vision-language models.
This repository is a mirror. The citable source of record is the deposit in Dane
Badawcze UW (University of Warsaw, ICM): https://doi.org/10.58132/FSJGSP, version 1.0.
The nine files in deposit/ are byte-identical to that release; deposit/schema.json
carries their sha256 sums so you can verify this yourself. Cite the DOI, not this URL.… See the full description on the dataset page: https://huggingface.co/datasets/msolarz/StatMapCorpus.
