datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EASI-Leaderboard-Data
EASI Leaderboard Data
A consolidated dataset for the EASI Leaderboard, containing the evaluation data (inputs/prompts) actually used on the leaderboard across spatial reasoning benchmarks for VLMs.
Looking for the Spatial Intelligence leaderboard?https://huggingface.co/spaces/lmms-lab-si/EASI-Leaderboard
🔎 Dataset Summary
Question types: MCQ (multiple choice) and NA (numeric answer).
File format: TSV only.
Usage: These TSVs are directly consumable by the EASI… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-si/EASI-Leaderboard-Data.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.MathVerse-lmmseval
Dataset Card for MathVerse
This is the version for lmms-eval. This shares the same data with the official dataset.
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Citation
Dataset Description
The capabilities of Multi-modal Large Language Models (MLLMs) in visual math problem-solving remain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual questions, which potentially… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MathVerse-lmmseval.EMMA
Dataset Description
EMMA (Enhanced MultiModal reAsoning) is a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding.
EMMA tasks demand advanced cross-modal reasoning that cannot be solved by thinking separately in each modality, offering an enhanced test suite for MLLMs' reasoning capabilities.
EMMA is composed of 2,788 problems, of which 1,796 are newly constructed, across four domains. Within each subject, we further provide… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/EMMA.PhysBench
PhysBench for lmms-eval
Normalized PhysBench annotations with the official public answers and media archives.
Source: https://huggingface.co/datasets/USC-PSI-Lab/PhysBench. The source dataset license is preserved.
EMMA-mini
Dataset Description
EMMA (Enhanced MultiModal reAsoning) is a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding.
EMMA tasks demand advanced cross-modal reasoning that cannot be solved by thinking separately in each modality, offering an enhanced test suite for MLLMs' reasoning capabilities.
EMMA is composed of 2,788 problems, of which 1,796 are newly constructed, across four domains. Within each subject, we further provide… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/EMMA-mini.PhysReason
PhysReason for lmms-eval
Normalized full and mini PhysReason configurations with Viewer-compatible image columns.
Source: https://huggingface.co/datasets/zhibei1204/PhysReason. The source dataset license is preserved.
