datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Video-MMELMMs-Eval-LiteegoschemaLVBenchNExTQAYouCook2LiveBenchhttps://arxiv.org/abs/2407.12772
TempCompassMMVP
MMVP (Multimodal Visual Patterns) Benchmark
This is a corrected version of the MMVP benchmark, re-hosted by lmms-lab-eval for use with lmms-eval.
Why this copy?
The original MMVP/MMVP dataset was uploaded in imagefolder format, which only exposes the image column. The text annotations (Question, Options, Correct Answer, Index) from the accompanying Questions.csv were not loaded into the dataset, making it unusable for evaluation.
This version reconstructs the complete… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/MMVP.ActivityNetQAVideoMMMUThis dataset contains the data for the paper Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos. Video-MMMU is a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos.
Project page: https://videommmu.github.io/
Leaderboard (last updated: 07 Feb, 2025)
Model
Overall
Perception
Comprehension
Adaptation
Δknowledge
Human Expert
74.44
84.33
78.67
60.33
+33.1… See the full description on the dataset page: https://huggingface.co/datasets/lmms-eval/VideoMMMU.MME-RealWorld-Lmms-evalMindCube_lmmseval
MindCube LMMs Eval Dataset
This dataset is formatted for use with lmms-eval framework.
Dataset Schema
Column
Type
Description
id
string
Unique identifier for each sample (format: {split}_{scene_id}_{question_id})
category
list[string]
Category labels (e.g., ["perpendicular", "P-O", "meanwhile", "self"])
type
string
Question type (e.g., "1_frame", "2_frame", "3_frame", "general")
meta_info
list[list[string]]
Metadata about scene objects and their spatial… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/MindCube_lmmseval.MathVerse-lmmseval
Dataset Card for MathVerse
This is the version for lmms-eval. This shares the same data with the official dataset.
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Citation
Dataset Description
The capabilities of Multi-modal Large Language Models (MLLMs) in visual math problem-solving remain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual questions, which potentially… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MathVerse-lmmseval.PerceptionTest_Valcharades_staVideoChatGPTvideo-tt
Towards Video Thinking Test (Video-TT): A Holistic Benchmark for Advanced Video Reasoning and Understanding
Video-TT comprises 1,000 YouTube videos, each paired with one open-ended question and four adversarial questions designed to probe visual and narrative complexity.
Paper: https://arxiv.org/abs/2507.15028
Project page: https://zhangyuanhan-ai.github.io/video-tt/
🚀 What's New
[2025.03] We release the benchmark!
1. Why Do We Need a New… See the full description on the dataset page: https://huggingface.co/datasets/lmms-eval/video-tt.Spatial457egotempo
EgoTempo
Full-set metadata for lmms-eval task egotempo.
Annotation source: https://raw.githubusercontent.com/google-research-datasets/egotempo/main/egotempo_openQA.json
Raw video location: Ego4D clips (license-gated), clip id in clip_id.
Expected local media root for evaluation: $EGOTEMPO_VIDEO_DIR or $HF_HOME/egotempo.
TOMATOViewSpatial_lmmseval3DSRBench_lmmseval
3DSRBench (lmms-eval compatible)
This is a reformatted version of 3DSRBench for compatibility with lmms-eval.
Dataset Description
3DSRBench is a comprehensive 3D spatial reasoning benchmark that evaluates the 3D spatial reasoning capabilities of Large Multimodal Models (LMMs). It includes 2,100 VQAs on MS-COCO images and 672 on multi-view synthetic images rendered from HSSD.
Subsets
This dataset contains two subsets:
1. 3dsr… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/3DSRBench_lmmseval.MME-RealWorld-lite-lmms-eval
2024.11.14 🌟 MME-RealWorld now has a lite version (50 samples per task, or all if fewer than 50) for inference acceleration, which is also supported by VLMEvalKit and Lmms-eval.
2024.09.03 🌟 MME-RealWorld is now supported in the VLMEvalKit and Lmms-eval repository, enabling one-click evaluation—give it a try!"
2024.08.20 🌟 We are very proud to launch MME-RealWorld, which contains 13K high-quality images, annotated by 32 volunteers, resulting in 29K question-answer pairs that cover 43… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/MME-RealWorld-lite-lmms-eval.MME-RealWorld-CN-Lmms-evalPerceptionTestMMVUHLE-Verified
HLE-Verified (HF-native)
This dataset is a Hugging Face-native conversion of skylenage/HLE-Verified at revision becad9f339dfce27df0ebb38e55dabef12ca5735.
Why this exists
The source dataset stores nested verification fields with mixed runtime types (for example 0/1/"uncertain"), which breaks strict Arrow JSON parsing in datasets.load_dataset.
This converted dataset normalizes those fields and publishes split-ready JSONL files for direct use in lmms_eval.
Split… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/HLE-Verified.VATEXGoogleDeepMind-NEPTUNE
