datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VLM-SubtleBench
VLM-SubtleBench
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/VLM-SubtleBench.vlmsareblindArXiv - Website
vlms-are-biased
Vision Language Models are Biased
by
An Vo1*,
Khai-Nguyen Nguyen2*,
Mohammad Reza Taesiri3,
Vy Tuong Dang1,
Anh Totti Nguyen4†,
Daeyoung Kim1†
*Equal contribution †Equal advising
1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University
TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.opht_vlms_lowopht_vlms_low_selectedvlms-are-biased
Vision Language Models are Biased
by
An Vo1*,
Khai-Nguyen Nguyen2*,
Mohammad Reza Taesiri3,
Vy Tuong Dang1,
Anh Totti Nguyen4†,
Daeyoung Kim1†
*Equal contribution †Equal advising
1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University
TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vlms-are-biased.VWSD-VLMsThis is our dataset obtained by exploiting the co-Hyphonym relation in BabelNet for VWSD.
We provide train, test and validation splits. Test and validation have 10 candidate images, while in the train set there are fewer candidates.
We also provide the dataset already formatted for generative VLM fine-tuning and evaluation.
Finally, we also provide the images associated with the dataset here.
VLM-SFTvlms-are-confused-tourists
Vision Language Models are Confused Tourists ✈️ 🤔
[!NOTE]
We are still in the process of beautifying the README of the HF dataset. Nonetheless, our data is fully usable!
Although the cultural dimension has been one of the key aspects in evaluating Vision-Language Models (VLMs), their ability to remain stable across diverse cultural inputs remains largely untested, despite being crucial to support diversity and multicultural societies. Existing evaluations often rely on… See the full description on the dataset page: https://huggingface.co/datasets/patrickamadeus/vlms-are-confused-tourists.vlm-safety-inspector-dataset
Safety Inspector V2 LoRA Training Dataset & Hyperparameter Specification
This dataset repository contains the offline warm-up SFT dataset and standardized LoRA training configuration for training the Vision-Language Model (VLM) Safety Inspector on the 50 tabletop manipulation scenes (Split into 45 Train + 5 Validation).
1. Dataset Overview
Source Scenes: 50 Tabletop Scenes (45 Train, 5 Val, 0 Test)
Task Levels: single_step (1 primitive), safety_two_step (2… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/vlm-safety-inspector-dataset.vlm_shape_benchmarkvlm-segmentation-cotvlm-segmentation-cotWhatsUp_VLMsHard_images_for_VLMsVLM_semantics_SLO_benchmark
VLM Semantics SLO Benchmark
VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia.
The released JSON contains 4,950 image-level records. Every record has ten… See the full description on the dataset page: https://huggingface.co/datasets/maticmatusek/VLM_semantics_SLO_benchmark.vlm-sft-mix-en-ko-26-06
VLM SFT Mix — English + Korean (26-06)
A unified multimodal supervised fine-tuning (SFT) mix for vision-language models (Gemma / LLaVA family). It combines vision (image–conversation) and text (instruction / reasoning) data in parallel English and Korean. Images are embedded as bytes inside the parquet files, so the dataset loads directly with 🤗 datasets — no separate image files to download.
Summary
Configs (sub-datasets)
175
Examples (rows)
42… See the full description on the dataset page: https://huggingface.co/datasets/Yong-Hoon/vlm-sft-mix-en-ko-26-06.VLMsAreBlindvlm_splitstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetVLM_SingleAction2whatsup_vlmsstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetOPD_vlmsarebiased
Dataset Card for "OPD_vlmsarebiased"
More Information needed
structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetvlm-sqr-3vlm-sqrstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetvlm-sqr-2project_x
