CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KRAFTON /VLM-SubtleBench VLM-SubtleBench VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning? The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/VLM-SubtleBench.imagevisual-question-answering10K<n<100K5 likes2.8k downloads7mo agoHugging Face02XAI /vlmsareblindArXiv - Website imagequestion-answering1K<n<10K28 likes2k downloads2y agoHugging Face03anvo25 /vlms-are-biased Vision Language Models are Biased by An Vo1*, Khai-Nguyen Nguyen2*, Mohammad Reza Taesiri3, Vy Tuong Dang1, Anh Totti Nguyen4†, Daeyoung Kim1† *Equal contribution    †Equal advising 1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.imagevisual-question-answering10K<n<100K29 likes1.4k downloads10mo agoHugging Face04LangAGI-Lab /opht_vlms_lowimage10K<n<100K0 likes1k downloads1y agoHugging Face05LangAGI-Lab /opht_vlms_low_selectedimage1K<n<10K0 likes526 downloads1y agoHugging Face06Rendra86318 /vlms-are-biased Vision Language Models are Biased by An Vo1*, Khai-Nguyen Nguyen2*, Mohammad Reza Taesiri3, Vy Tuong Dang1, Anh Totti Nguyen4†, Daeyoung Kim1† *Equal contribution    †Equal advising 1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vlms-are-biased.imagevisual-question-answering10K<n<100K0 likes98 downloads9mo agoHugging Face07swap-uniba /VWSD-VLMsThis is our dataset obtained by exploiting the co-Hyphonym relation in BabelNet for VWSD. We provide train, test and validation splits. Test and validation have 10 candidate images, while in the train set there are fewer candidates. We also provide the dataset already formatted for generative VLM fine-tuning and evaluation. Finally, we also provide the images associated with the dataset here. image0 likes92 downloads1y agoHugging Face08YangyiYY /VLM-SFTimagetext-generation1M<n<10M2 likes87 downloads2y agoHugging Face09patrickamadeus /vlms-are-confused-tourists Vision Language Models are Confused Tourists ✈️ 🤔 [!NOTE] We are still in the process of beautifying the README of the HF dataset. Nonetheless, our data is fully usable! Although the cultural dimension has been one of the key aspects in evaluating Vision-Language Models (VLMs), their ability to remain stable across diverse cultural inputs remains largely untested, despite being crucial to support diversity and multicultural societies. Existing evaluations often rely on… See the full description on the dataset page: https://huggingface.co/datasets/patrickamadeus/vlms-are-confused-tourists.imagevisual-question-answering1K<n<10K3 likes79 downloads9mo agoHugging Face10lvesucces /vlm-safety-inspector-dataset Safety Inspector V2 LoRA Training Dataset & Hyperparameter Specification This dataset repository contains the offline warm-up SFT dataset and standardized LoRA training configuration for training the Vision-Language Model (VLM) Safety Inspector on the 50 tabletop manipulation scenes (Split into 45 Train + 5 Validation). 1. Dataset Overview Source Scenes: 50 Tabletop Scenes (45 Train, 5 Val, 0 Test) Task Levels: single_step (1 primitive), safety_two_step (2… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/vlm-safety-inspector-dataset.imagevisual-question-answering1K<n<10K0 likes66 downloads10d agoHugging Face11ajaymin28 /vlm_shape_benchmarkimage1K<n<10K0 likes65 downloads1y agoHugging Face12huyhoangdinhcong /vlm-segmentation-cotimage10K<n<100K1 likes60 downloads1y agoHugging Face13VietMedTeam /vlm-segmentation-cotimage10K<n<100K0 likes52 downloads9mo agoHugging Face14ohjoonhee /WhatsUp_VLMsimage1K<n<10K0 likes48 downloads1y agoHugging Face15MichielBontenbal /Hard_images_for_VLMsimagequestion-answeringn<1K0 likes39 downloads2y agoHugging Face16maticmatusek /VLM_semantics_SLO_benchmark VLM Semantics SLO Benchmark VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia. The released JSON contains 4,950 image-level records. Every record has ten… See the full description on the dataset page: https://huggingface.co/datasets/maticmatusek/VLM_semantics_SLO_benchmark.imagevisual-question-answering1K<n<10K0 likes33 downloads3d agoHugging Face17Yong-Hoon /vlm-sft-mix-en-ko-26-06gated VLM SFT Mix — English + Korean (26-06) A unified multimodal supervised fine-tuning (SFT) mix for vision-language models (Gemma / LLaVA family). It combines vision (image–conversation) and text (instruction / reasoning) data in parallel English and Korean. Images are embedded as bytes inside the parquet files, so the dataset loads directly with 🤗 datasets — no separate image files to download. Summary Configs (sub-datasets) 175 Examples (rows) 42… See the full description on the dataset page: https://huggingface.co/datasets/Yong-Hoon/vlm-sft-mix-en-ko-26-06.image10M<n<100M0 likes31 downloads20d agoHugging Face18mm-eval /VLMsAreBlindimage1K<n<10K0 likes28 downloads2mo agoHugging Face19flipwooyoung /vlm_splitimage1K<n<10K0 likes27 downloads2y agoHugging Face20davidberenstein1957 /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes13 downloads2y agoHugging Face21zesquirrelnator /VLM_SingleAction2imagen<1K0 likes12 downloads2y agoHugging Face22amitakamath2 /whatsup_vlmsimagen<1K0 likes12 downloads1y agoHugging Face23Hoang123 /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes12 downloads1y agoHugging Face24tomyoon2 /OPD_vlmsarebiased Dataset Card for "OPD_vlmsarebiased" More Information needed image1K<n<10K0 likes12 downloads8mo agoHugging Face25paulwoodward /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes10 downloads2y agoHugging Face26orcn /vlm-sqr-3imagen<1K0 likes9 downloads2y agoHugging Face27orcn /vlm-sqrimagen<1K0 likes8 downloads2y agoHugging Face28chandc /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes8 downloads1y agoHugging Face29orcn /vlm-sqr-2imagen<1K0 likes6 downloads2y agoHugging Face30vlmsarebiased /project_ximage10K<n<100K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.