CoolFace
Datasetpublic

DickMan42/MMVP

MMVP Benchmark Datacard Basic Information Title: MMVP Benchmark Description: The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying “CLIP-blind pairs” – images that are perceived as similar by CLIP despite having clear visual differences. MMVP benchmarks the performance of state-of-the-art systems, including GPT-4V, across nine basic visual patterns. It highlights the challenges these systems face in answering straightforward questions, often… See the full description on the dataset page: https://huggingface.co/datasets/DickMan42/MMVP.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes10downloads
Dataset Card

MMVP Benchmark Datacard

Basic Information

Title: MMVP Benchmark

Description: The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying “CLIP-blind pairs” – images that are perceived as similar by CLIP despite having clear visual differences. MMVP benchmarks the performance of state-of-the-art systems, including GPT-4V, across nine basic visual patterns. It highlights the challenges these systems face in answering straightforward questions, often leading to incorrect responses and hallucinated explanations.

Dataset Details

  • Content Types: Images (CLIP-blind pairs)
  • Volume: 300 images
  • Source of Data: Derived from ImageNet-1k and LAION-Aesthetics
  • Data Collection Method: Identification of CLIP-blind pairs through comparative analysis