VLMs
Datasets
All datasets matching “VLMs”VLM-SubtleBench
VLM-SubtleBench
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/VLM-SubtleBench.vlmsareblindArXiv - Website
vlms-are-biased
Vision Language Models are Biased
by
An Vo1*,
Khai-Nguyen Nguyen2*,
Mohammad Reza Taesiri3,
Vy Tuong Dang1,
Anh Totti Nguyen4†,
Daeyoung Kim1†
*Equal contribution †Equal advising
1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University
TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.opht_vlms_lowopht_vlms_low_selectedvlm-safety-circuits
VLMSafe-420: A Multimodal Safety Benchmark for VLM Circuit Analysis
A 420-entry multimodal safety dataset with counterfactual pairs, designed for mechanistic interpretability of safety circuits in Vision-Language Models.
Dataset Description
Each entry contains a harmful/benign counterfactual pair spanning 38 safety categories including 50 JailbreakBench-style prompts. The dataset covers three counterfactual types:
Type
Count
Description
Text… See the full description on the dataset page: https://huggingface.co/datasets/ArthT/vlm-safety-circuits.
