datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMVU
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
🌐 Homepage •
🥇 Leaderboard •
📖 Paper •
🤗 Data
📰 News
2025-01-21: We are excited to release the MMVU paper, dataset, and evaluation code!
👋 Overview
Why MMVU Benchmark?
Despite the rapid progress of foundation models in both text-based and image-based expert reasoning, there is a clear gap in evaluating these models’ capabilities in specialized-domain video understanding.… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/MMVU.MMVP
MMVP (Multimodal Visual Patterns) Benchmark
This is a corrected version of the MMVP benchmark, re-hosted by lmms-lab-eval for use with lmms-eval.
Why this copy?
The original MMVP/MMVP dataset was uploaded in imagefolder format, which only exposes the image column. The text annotations (Question, Options, Correct Answer, Index) from the accompanying Questions.csv were not loaded into the dataset, making it unusable for evaluation.
This version reconstructs the complete… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/MMVP.MMVet
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of MM-Vet. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@misc{yu2023mmvet,
title={MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/MMVet.mm-vetPaper: https://arxiv.org/abs/2308.02490
MMVUmm_vqav2mcot_r1_mcq_66kmm_visual7wmm-vet-v2Paper: https://arxiv.org/abs/2408.00765
MMVetmcot_r1_vqa_66kMMV-dataset
MMV-dataset
Reproducible research splits and full-training subsets derived from long-video benchmarks. No video, audio, or subtitle media is included.
Hosted data
Component
Train
Validation
Grouping unit
License
LongVideoBench labeled validation set
401
936
video_id
CC BY-NC-SA 4.0
EgoTempo open-ended QA
150
350
original Ego4D video UID
CC BY 4.0
Unified supervised records
1,084
2,529
inherited
mixed; see per-row license
The LongVideoBench… See the full description on the dataset page: https://huggingface.co/datasets/czty/MMV-dataset.x2x_rft_22kx2x_rft_16kMMVU_MCMMVMBench_VQAMMVPMMVU-VQA
MMVUVideoCentricQA
An MTEB dataset
Massive Text Embedding Benchmark
MMVU is an expert-level, multi-discipline video understanding benchmark with questions spanning 27 subjects across Science, Healthcare, Humanities & Social Sciences, and Engineering. Each multiple-choice example pairs a specialized-domain video with a question and 5 candidate answers. The task is formulated as multiple-choice retrieval: given the (video, question) pair, retrieve the correct candidate. Used the public… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MMVU-VQA.mm_vistextmm_visualmrcMMVUMM_VET_modif
Dataset Card for "MM_VET_modif"
MM-VET Benchmark
MM-Verify-Datamm_vqaradMM-Vet-v2
⚠️ DEPRECATED
This repository is deprecated — use mm-eval/MMVet-v2 instead.
mm-eval/MM-Vet-v2 and mm-eval/MMVet-v2 are duplicate uploads of the same
benchmark (MM-Vet v2, 517 identical rows — same ids, questions, and answers;
confirmed in the 2026-07-07 org audit). Per the owner's decision the newer
conversion MMVet-v2 is the canonical copy. The data here is kept unchanged
for reproducibility of past runs; do not use it for new evaluations.
MMVU-VQA
MMVUVideoCentricQA
An MTEB dataset
Massive Text Embedding Benchmark
MMVU is an expert-level, multi-discipline video understanding benchmark with questions spanning 27 subjects across Science, Healthcare, Humanities & Social Sciences, and Engineering. Each multiple-choice example pairs a specialized-domain video with a question and 5 candidate answers. The task is formulated as multiple-choice retrieval: given the (video, question) pair, retrieve the correct candidate. Used the public… See the full description on the dataset page: https://huggingface.co/datasets/Wissam42/MMVU-VQA.MMVet-v2MMVetMMVP
MMVP Benchmark
refactor MMVP to support VLMEalKit
Benchmark Information
number of questions: 300
question type: multiple choice question
question format: image + text
Reference
VLMEvalKit
MMVP
MMVet
