CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KRAFTON /VLM-SubtleBench VLM-SubtleBench VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning? The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/VLM-SubtleBench.imagevisual-question-answering10K<n<100K5 likes2.8k downloads7mo agoHugging Face02XAI /vlmsareblindArXiv - Website imagequestion-answering1K<n<10K28 likes1.9k downloads2y agoHugging Face03anvo25 /vlms-are-biased Vision Language Models are Biased by An Vo1*, Khai-Nguyen Nguyen2*, Mohammad Reza Taesiri3, Vy Tuong Dang1, Anh Totti Nguyen4†, Daeyoung Kim1† *Equal contribution    †Equal advising 1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.imagevisual-question-answering10K<n<100K29 likes1.2k downloads10mo agoHugging Face04LangAGI-Lab /opht_vlms_lowimage10K<n<100K0 likes1.1k downloads1y agoHugging Face05LangAGI-Lab /opht_vlms_low_selectedimage1K<n<10K0 likes530 downloads1y agoHugging Face06swap-uniba /VWSD-VLMsThis is our dataset obtained by exploiting the co-Hyphonym relation in BabelNet for VWSD. We provide train, test and validation splits. Test and validation have 10 candidate images, while in the train set there are fewer candidates. We also provide the dataset already formatted for generative VLM fine-tuning and evaluation. Finally, we also provide the images associated with the dataset here. image0 likes92 downloads1y agoHugging Face07Rendra86318 /vlms-are-biased Vision Language Models are Biased by An Vo1*, Khai-Nguyen Nguyen2*, Mohammad Reza Taesiri3, Vy Tuong Dang1, Anh Totti Nguyen4†, Daeyoung Kim1† *Equal contribution    †Equal advising 1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vlms-are-biased.imagevisual-question-answering10K<n<100K0 likes90 downloads9mo agoHugging Face08lvesucces /vlm-safety-inspector-dataset Safety Inspector V2 LoRA Training Dataset & Hyperparameter Specification This dataset repository contains the offline warm-up SFT dataset and standardized LoRA training configuration for training the Vision-Language Model (VLM) Safety Inspector on the 50 tabletop manipulation scenes (Split into 45 Train + 5 Validation). 1. Dataset Overview Source Scenes: 50 Tabletop Scenes (45 Train, 5 Val, 0 Test) Task Levels: single_step (1 primitive), safety_two_step (2… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/vlm-safety-inspector-dataset.imagevisual-question-answering1K<n<10K0 likes82 downloads11d agoHugging Face09patrickamadeus /vlms-are-confused-tourists Vision Language Models are Confused Tourists ✈️ 🤔 [!NOTE] We are still in the process of beautifying the README of the HF dataset. Nonetheless, our data is fully usable! Although the cultural dimension has been one of the key aspects in evaluating Vision-Language Models (VLMs), their ability to remain stable across diverse cultural inputs remains largely untested, despite being crucial to support diversity and multicultural societies. Existing evaluations often rely on… See the full description on the dataset page: https://huggingface.co/datasets/patrickamadeus/vlms-are-confused-tourists.imagevisual-question-answering1K<n<10K3 likes81 downloads9mo agoHugging Face10YangyiYY /VLM-SFTimagetext-generation1M<n<10M2 likes70 downloads2y agoHugging Face11huyhoangdinhcong /vlm-segmentation-cotimage10K<n<100K1 likes69 downloads1y agoHugging Face12ajaymin28 /vlm_shape_benchmarkimage1K<n<10K0 likes67 downloads1y agoHugging Face13VietMedTeam /vlm-segmentation-cotimage10K<n<100K0 likes61 downloads9mo agoHugging Face14ohjoonhee /WhatsUp_VLMsimage1K<n<10K0 likes60 downloads1y agoHugging Face15maticmatusek /VLM_semantics_SLO_benchmark VLM Semantics SLO Benchmark VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia. The released JSON contains 4,950 image-level records. Every record has ten… See the full description on the dataset page: https://huggingface.co/datasets/maticmatusek/VLM_semantics_SLO_benchmark.imagevisual-question-answering1K<n<10K0 likes44 downloads4d agoHugging Face16MichielBontenbal /Hard_images_for_VLMsimagequestion-answeringn<1K0 likes32 downloads2y agoHugging Face17Yong-Hoon /vlm-sft-mix-en-ko-26-06gated VLM SFT Mix — English + Korean (26-06) A unified multimodal supervised fine-tuning (SFT) mix for vision-language models (Gemma / LLaVA family). It combines vision (image–conversation) and text (instruction / reasoning) data in parallel English and Korean. Images are embedded as bytes inside the parquet files, so the dataset loads directly with 🤗 datasets — no separate image files to download. Summary Configs (sub-datasets) 175 Examples (rows) 42… See the full description on the dataset page: https://huggingface.co/datasets/Yong-Hoon/vlm-sft-mix-en-ko-26-06.image10M<n<100M0 likes32 downloads22d agoHugging Face18flipwooyoung /vlm_splitimage1K<n<10K0 likes29 downloads2y agoHugging Face19mm-eval /VLMsAreBlindimage1K<n<10K0 likes29 downloads2mo agoHugging Face20davidberenstein1957 /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes14 downloads2y agoHugging Face21Hoang123 /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes14 downloads1y agoHugging Face22tomyoon2 /OPD_vlmsarebiased Dataset Card for "OPD_vlmsarebiased" More Information needed image1K<n<10K0 likes14 downloads8mo agoHugging Face23amitakamath2 /whatsup_vlmsimagen<1K0 likes13 downloads2y agoHugging Face24zesquirrelnator /VLM_SingleAction2imagen<1K0 likes10 downloads2y agoHugging Face25paulwoodward /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes10 downloads2y agoHugging Face26chandc /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes10 downloads1y agoHugging Face27orcn /vlm-sqr-3imagen<1K0 likes9 downloads2y agoHugging Face28orcn /vlm-sqrimagen<1K0 likes8 downloads2y agoHugging Face29vlmsarebiased /project_ximage10K<n<100K0 likes8 downloads1y agoHugging Face30orcn /vlm-sqr-2imagen<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.