mmbench
Datasets
All datasets matching “mmbench”MMBenchMMBench-ru
MMBench-ru
This is a translated version of original MMBench dataset and
stored in format supported for lmms-eval pipeline.
For this dataset, we:
Translate the original one with gpt-4o
Filter out unsuccessful translations, i.e. where the model protection was triggered
Manually validate most common errors
Dataset Structure
Dataset includes only dev split that is translated from dev split in lmms-lab/MMBench_EN.
Dataset contains 3910 samples in the same to… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/MMBench-ru.mmbench2
MMBench2
Hallucination in World Models is Predictable and Preventable
Nicklas Hansen · Xiaolong Wang · UC San Diego
MMBench2 is a large-scale dataset for visual world modeling, accompanying the paper Hallucination in World Models is Predictable and Preventable. It spans 210 continuous control tasks across 10 domains (DMControl, DMControl Extended, Meta-World, ManiSkill3, MuJoCo, MiniArcade, Box2D, RoboDesk, OGBench, and Atari) comprising 65,600… See the full description on the dataset page: https://huggingface.co/datasets/nicklashansen/mmbench2.MM-SpuBench
MM-SpuBench Datacard
Basic Information
Title: The Multimodal Spurious Benchmark (MM-SpuBench)
Description: MM-SpuBench is a comprehensive benchmark designed to evaluate the robustness of MLLMs to spurious biases. This benchmark systematically assesses how well these models distinguish between core and spurious features, providing a detailed framework for understanding and quantifying spurious biases.
Data Structure:
├── data/images
│ ├── 000000.jpg
│ ├── 000001.jpg
│… See the full description on the dataset page: https://huggingface.co/datasets/mmbench/MM-SpuBench.MMBench_EN
Dataset Card for "MMBench_EN"
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of the English subset of MMBench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{MMBench,
author = {Yuan Liu, Haodong… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/MMBench_EN.MMBench
Dataset Card for "MMBench"
More Information needed
