DickMan42/MMVP
MMVP Benchmark Datacard Basic Information Title: MMVP Benchmark Description: The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying “CLIP-blind pairs” – images that are perceived as similar by CLIP despite having clear visual differences. MMVP benchmarks the performance of state-of-the-art systems, including GPT-4V, across nine basic visual patterns. It highlights the challenges these systems face in answering straightforward questions, often… See the full description on the dataset page: https://huggingface.co/datasets/DickMan42/MMVP.
010
1---2license: mit3task_categories:4- question-answering5size_categories:6- n<1K7---8 9# MMVP Benchmark Datacard10 11## Basic Information12 13**Title:** MMVP Benchmark14 15**Description:** The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying “CLIP-blind pairs” – images that are perceived as similar by CLIP despite having clear visual differences. MMVP benchmarks the performance of state-of-the-art systems, including GPT-4V, across nine basic visual patterns. It highlights the challenges these systems face in answering straightforward questions, often leading to incorrect responses and hallucinated explanations.16 17## Dataset Details18 19- **Content Types:** Images (CLIP-blind pairs)20- **Volume:** 300 images21- **Source of Data:** Derived from ImageNet-1k and LAION-Aesthetics22- **Data Collection Method:** Identification of CLIP-blind pairs through comparative analysis23 