datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VLM4Bio
Dataset Card for VLM4Bio
Instructions for downloading the dataset
Install Git LFS
Git clone the VLM4Bio repository to download all metadata and associated files
Run the following commands in a terminal:
git clone https://huggingface.co/datasets/imageomics/VLM4Bio
cd VLM4Bio
Downloading and processing bird images
To download the bird images, run the following command:
bash download_bird_images.sh
This should download the bird images inside datasets/Bird/images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/VLM4Bio.OmniMed_VLMllava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.RefCOCO-VLMEvalKitSEEDBench2
Dataset Card for Dataset Name
The SEEDBench2 evaluation dataset hosted by VLMEval (authorized by the author).
Dataset Details
Language(s) (NLP): English
License: Apache 2.0
Repository: https://github.com/AILab-CVC/SEED-Bench
Paper [optional]: https://arxiv.org/abs/2311.17092
Citation
@misc{li2023seedbench2,
title={SEED-Bench-2: Benchmarking Multimodal Large Language Models},
author={Bohao Li and Yuying Ge and Yixiao Ge and Guangzhi Wang and Rui… See the full description on the dataset page: https://huggingface.co/datasets/VLMEval/SEEDBench2.GMAI-MMBenchTallyQA-VLMEvalKitOmni3DBench-VLMEvalKitVLMEvalKit_CVQA
CVQA for VLMEvalKit
Original dataset: ported to VLMEvalKit
From the original authors:
CVQA is a culturally diverse multilingual VQA benchmark consisting of over 10,000 questions from 39 country-language pairs. The questions in CVQA are written in both the native languages and English, and are categorized into 10 diverse categories.
{'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=2048x1536 at 0x7C3E0EBEEE00>,
'ID': '5919991144272485961_0',
'Subset':… See the full description on the dataset page: https://huggingface.co/datasets/timothycdc/VLMEvalKit_CVQA.vlmevalkit_filesViewSpatial-Bench-vlmevalMathverse_VLMEvalKitMedical_VLM_SycophancyThis the official data hosting repository for paper "EchoBench: Benchmarking Sycophancy in Medical
Large Vision Language Models".
============open-source_models============
For experiments on open-source models, our implementation is built upon the VLMEvalkit framework.
Navigate to the VLMEval directory
Set up the environment by running: "pip install -e ."
Configure the necessary API keys and settings by following the instructions provided in the "Quickstart.md" file of VLMEvalkit.
To… See the full description on the dataset page: https://huggingface.co/datasets/Botai666/Medical_VLM_Sycophancy.LiveXiv-VLMEvalKitcrpe_vlmevalkitwilddoc-vlmevalGSM8K-V-VLMEvalKitMME-CoT_VLMEvalKitHateful_Memes_in_VLMThis dataset contains the response of VLMs (InstructBlip, ShareGPT4V, LLaVA and CogVLM) to hateful memes and the annotation to these responses. For more information, please refer to paper "From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Models."
chart_cap_vlmevalkitvlmevalkit-refcocogvlm-eval-videos
VLM Eval Videos
A video benchmark dataset for evaluating Vision–Language Models (VLMs) on short-form action recognition.
Each clip is paired with a fixed question and a ground-truth short-sentence answer,
making it suitable for automated VLM inference pipelines and LLM-as-a-judge scoring.
Dataset Details
Description
VLM Eval Videos contains 693 short MP4 video clips drawn from YouTube, organised into five categories.
Four categories contain clips of… See the full description on the dataset page: https://huggingface.co/datasets/gnitoahc/vlm-eval-videos.CountBenchQA-VLMEvalKitVLMEvalKit_AyaVisionBench
Aya Vision Bench for VLMEvalKit
Original dataset: ported to VLMEvalKit
Multilingual dataset spans 23 languages and 9 distinct task categories, with 15 samples per category, resulting in 135 image-question pairs per language.
Original dataset row:
{'image': [PIL.Image],
'image_source': 'VisText',
'image_source_category': 'Chart/figure understanding',
'index' : '17'
'question': 'If the top three parties by vote percentage formed a coalition, what percentage of the total votes… See the full description on the dataset page: https://huggingface.co/datasets/timothycdc/VLMEvalKit_AyaVisionBench.MicroVQA_VLMtab-vlm
TAB-VLM: Temporal Anachronism Benchmark for Vision-Language Models
Paper: On the Cultural Anachronism and Temporal Reasoning in Vision Language Models (ACL 2026 Findings)
Authors: Mukul Ranjan, Prince Jha, Khushboo Kumari, Zhiqiang Shen
TAB-VLM is a benchmark for measuring cultural anachronism in Vision-Language Models — the tendency to misinterpret historical artifacts using temporally inappropriate concepts, materials, or cultural frameworks. The benchmark consists of 600… See the full description on the dataset page: https://huggingface.co/datasets/mukul54/tab-vlm.Vie-ScenicDesign2Code-VLMEvalKitOmniMat1K-VLMEvalKit
OmniMatBench 1K Subset for VLMEvalKit
Scope notice: This repository contains OmniMat1K, a 1,000-item
subset with 498 QA and 502 CAL records. It is not the complete
3,171-item OmniMatBench release described in the paper. Evaluation results
produced from this repository must be labeled OmniMat1K and must not be
presented as scores on the complete benchmark.
Release status
This repository contains the 1,000-item OmniMatBench subset prepared for
VLMEvalKit. The… See the full description on the dataset page: https://huggingface.co/datasets/Summer12138/OmniMat1K-VLMEvalKit.VLMEvalKit_sourced_DocVQA_val
