TIGER-Lab/Mantis-Eval
Overview This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/. Statistics This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems. Leaderboard We list the current results as follows: Models Size Mantis-Eval LLaVA OneVision 72B 77.60 LLaVA OneVision 7B 64.20 GPT-4V -… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Eval.
Overview
This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/.
Statistics
This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems.
Leaderboard
We list the current results as follows:
Citation
If you are using this dataset, please cite our work with
@article{Jiang2024MANTISIM,
title={MANTIS: Interleaved Multi-Image Instruction Tuning},
author={Dongfu Jiang and Xuan He and Huaye Zeng and Cong Wei and Max W.F. Ku and Qian Liu and Wenhu Chen},
journal={Transactions on Machine Learning Research},
year={2024},
volume={2024},
url={https://openreview.net/forum?id=skLtdUVaJa}
}