datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.PerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.perception-sets-matrix
Perception Sets the Matrix — Match-Propensity Campaign
Complete, reproducible campaign data for the pre-registered test of the
spectral-match linking model: at cohort level, the perceptual match
between an observer cohort's eight-dimensional perception profile and a
brand's profile is a monotone predictor of the cohort's stated choice
propensity within its category consideration set — the generative layer
under the stochastic brand-choice tradition's exogenous switching… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/perception-sets-matrix.PerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/Bithubs/PerceptionBench.Perception-Collection
Dataset Card
Homepage: https://kaistai.github.io/prometheus-vision/
Repository: https://github.com/kaistAI/prometheus-vision
Paper: https://arxiv.org/abs/2401.06591
Point of Contact: seongyun@kaist.ac.kr
Dataset summary
Perception Collection is the first multi-modal feedback dataset that could be used to train an evaluator VLM. Perception Collection includes 15K fine-grained criteria that determine the crucial aspect for each instance.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Perception-Collection.Perception-Bench
Dataset Card
Homepage: https://kaistai.github.io/prometheus-vision/
Repository: https://github.com/kaistAI/prometheus-vision
Paper: https://arxiv.org/abs/2401.06591
Point of Contact: seongyun@kaist.ac.kr
Dataset summary
Perception-Bench is a benchmark for evaluating the long-form response of a VLM (Vision Language Model) across various domains of images, and it is a held-out test
set of the Perception-Collection
Languages
English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Perception-Bench.Synthetic-Cyclic-Perception_exp1
'Make knowledge free for everyone'
Synthetic-Cyclic-Perception
Synthetic-Cyclic-Perception is a synthetic visual dataset created by iteratively generating and describing images in a cyclic fashion. Each cycle begins with a seed prompt that feeds into a diffusion model to generate an initial image. A vision model then provides a detailed description of the generated image, which becomes the prompt for the next cycle. This process is repeated across multiple cycles and batches… See the full description on the dataset page: https://huggingface.co/datasets/DevQuasar/Synthetic-Cyclic-Perception_exp1.VLM-CapCurriculum-Perception-Data
VLM-CapCurriculum-Perception (D_perc)
Stage-1 visual perception data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
Each sample is a 4-way multiple-choice question over an image where the question can be answered from a fine-grained image caption but is missed by a strong VLM looking only at the image — by construction, these samples isolate perception failures from… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-Perception-Data.robotics_perceptionRobotics
han-environment-perception-v1
Humanoid Environment Perception Dataset
This dataset represents how humanoid robots perceive
human-centered environments before making decisions.
The data focuses on abstract perception rather than
raw sensor values, making it hardware-independent.
Contains
Environment description
Visual observations
Audio cues
Contextual role
Part of
Humanoid Network (HAN)
License
MIT
han-humanoid-perception-anomaly-signals-v1
Humanoid Perception Anomaly Signals (HPAS)
Abstract
HPAS contains structured anomaly signals detected
from humanoid sensor fusion systems during operation.
The dataset supports anomaly detection, sensor
confidence modeling, and environmental uncertainty research.
Fields
timestamp
sensor_type
raw_signal_variance
anomaly_score
confidence_level
environment_context
Intended Research
Sensor anomaly detection
Confidence calibration
Environmental… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-perception-anomaly-signals-v1.han-hospital-perception-v1
Hospital Environment Perception Dataset
This dataset represents how humanoid robots perceive
hospital and care environments in a safe, human-centric way.
The focus is on abstract perception to support
privacy-preserving and hardware-independent systems.
Contains
Care environment description
Visual observations
Audio cues
Robot assistance role
Part of
Humanoid Network (HAN)
License
MIT
humanoid-multi-sensor-perception-dataset
Humanoid Multi-Sensor Perception Dataset
Dataset for training humanoid systems to interpret
multi-sensor inputs including vision, audio, and proximity signals.
Description
Provides synchronized sensor readings for real-time
environment perception and context awareness.
File
multi_sensor_perception_dataset.json
License
MIT
han-security-perception-v1
Security Environment Perception Dataset
This dataset represents how humanoid robots perceive
security-sensitive environments in a non-intrusive,
human-aware manner.
The dataset focuses on situational awareness
rather than threat or force.
Contains
Secured environment description
Visual observations
Audio cues
Monitoring role context
Part of
Humanoid Network (HAN)
License
MIT
robotics_perception_textdriving_perception.jsonhan-warehouse-perception-v1
Warehouse Environment Perception Dataset
This dataset represents how humanoid robots perceive
industrial warehouse environments in a human-centric way.
The data focuses on abstract observations rather than
raw sensor data, enabling cross-robot compatibility.
Contains
Warehouse environment description
Visual observations
Audio cues
Robot operational role
Part of
Humanoid Network (HAN)
License
MIT
humanoid-perception-dataset-v2uncertainty-perception
