datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LMMs-Eval-LiteLiveBenchhttps://arxiv.org/abs/2407.12772
MMVP
MMVP (Multimodal Visual Patterns) Benchmark
This is a corrected version of the MMVP benchmark, re-hosted by lmms-lab-eval for use with lmms-eval.
Why this copy?
The original MMVP/MMVP dataset was uploaded in imagefolder format, which only exposes the image column. The text annotations (Question, Options, Correct Answer, Index) from the accompanying Questions.csv were not loaded into the dataset, making it unusable for evaluation.
This version reconstructs the complete… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/MMVP.VideoMMMUThis dataset contains the data for the paper Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos. Video-MMMU is a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos.
Project page: https://videommmu.github.io/
Leaderboard (last updated: 07 Feb, 2025)
Model
Overall
Perception
Comprehension
Adaptation
Δknowledge
Human Expert
74.44
84.33
78.67
60.33
+33.1… See the full description on the dataset page: https://huggingface.co/datasets/lmms-eval/VideoMMMU.MindCube_lmmseval
MindCube LMMs Eval Dataset
This dataset is formatted for use with lmms-eval framework.
Dataset Schema
Column
Type
Description
id
string
Unique identifier for each sample (format: {split}_{scene_id}_{question_id})
category
list[string]
Category labels (e.g., ["perpendicular", "P-O", "meanwhile", "self"])
type
string
Question type (e.g., "1_frame", "2_frame", "3_frame", "general")
meta_info
list[list[string]]
Metadata about scene objects and their spatial… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/MindCube_lmmseval.MathVerse-lmmseval
Dataset Card for MathVerse
This is the version for lmms-eval. This shares the same data with the official dataset.
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Citation
Dataset Description
The capabilities of Multi-modal Large Language Models (MLLMs) in visual math problem-solving remain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual questions, which potentially… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MathVerse-lmmseval.Spatial457ViewSpatial_lmmseval3DSRBench_lmmseval
3DSRBench (lmms-eval compatible)
This is a reformatted version of 3DSRBench for compatibility with lmms-eval.
Dataset Description
3DSRBench is a comprehensive 3D spatial reasoning benchmark that evaluates the 3D spatial reasoning capabilities of Large Multimodal Models (LMMs). It includes 2,100 VQAs on MS-COCO images and 672 on multi-view synthetic images rendered from HSSD.
Subsets
This dataset contains two subsets:
1. 3dsr… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/3DSRBench_lmmseval.lmms-eval-datasetPhysBench
PhysBench for lmms-eval
Normalized PhysBench annotations with the official public answers and media archives.
Source: https://huggingface.co/datasets/USC-PSI-Lab/PhysBench. The source dataset license is preserved.
NaturalBench-lmms-eval
(NeurIPS24) NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Baiqi Li1*, Zhiqiu Lin1*, Wenxuan Peng1*, Jean de Dieu Nyandwi1*, Daniel Jiang1, Zixian Ma2, Simran Khanuja1, Ranjay Krishna2†, Graham Neubig1†, Deva Ramanan1†.
1Carnegie Mellon University, 2University of Washington
Links:
| 🏠Home Page | 🤗HuggingFace | 🏆Leaderboard | 📖Paper |
Usage:
'''
This dataset consists of the content from naturalbench… See the full description on the dataset page: https://huggingface.co/datasets/BaiqiL/NaturalBench-lmms-eval.IllusionBenchAV_Odyssey_Bench_LMMs_EvalMMSI-Video-Bench_lmmseval
MMSI-Video-Bench
A video-based spatial intelligence benchmark for evaluating Multimodal Large Language Models (MLLMs).
Dataset Description
MMSI-Video-Bench tests models on:
Spatial reasoning
Motion understanding
Planning and prediction
Cross-video reasoning
Dataset Structure
MMSI-Video-Bench/
├── data/
│ └── test-00000-of-00001.parquet # 1106 samples
├── frames.zip # Extracted video frames
├── ref_images.zip #… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/MMSI-Video-Bench_lmmseval.PhysReason
PhysReason for lmms-eval
Normalized full and mini PhysReason configurations with Viewer-compatible image columns.
Source: https://huggingface.co/datasets/zhibei1204/PhysReason. The source dataset license is preserved.
RealUnifyvstar_bench_lmms_evalMindCube_lmms_evalLMMs-Eval-Litelmms-eval-blogUniMMMUlmms-eval-milebench
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/wangzexin2002/lmms-eval-milebench.lmms-eval-results
