wulin222/MME-Unify
2024.08.20 ๐ We are proud to open-source MME-Unify, a comprehensive evaluation framework designed to systematically assess U-MLLMs. Our Benchmark covers 10 tasks with 30 subtasks, ensuring consistent and fair comparisons across studies. Paper: https://arxiv.org/abs/2504.03641 Code: https://github.com/MME-Benchmarks/MME-Unify Project page: https://mme-unify.github.io/ How to use? You can download images in this repository and the final structure should look like this:โฆ See the full description on the dataset page: https://huggingface.co/datasets/wulin222/MME-Unify.
- `2024.08.20` ๐ We are proud to open-source MME-Unify, a comprehensive evaluation framework designed to systematically assess U-MLLMs. Our Benchmark covers 10 tasks with 30 subtasks, ensuring consistent and fair comparisons across studies.
Paper: https://arxiv.org/abs/2504.03641
Code: https://github.com/MME-Benchmarks/MME-Unify
Project page: https://mme-unify.github.io/
How to use?
You can download images in this repository and the final structure should look like this:
MME-Unify
โโโ CommonSense_Questions
โโโ Conditional_Image_to_Video_Generation
โโโ Fine-Grained_Image_Reconstruction
โโโ Math_Reasoning
โโโ Multiple_Images_and_Text_Interlaced
โโโ Single_Image_Perception_and_Understanding
โโโ Spot_Diff
โโโ Text-Image_Editing
โโโ Text-Image_Generation
โโโ Text-to-Video_Generation
โโโ Video_Perception_and_Understanding
โโโ Visual_CoTDataset details
We present MME-Unify, a comprehensive evaluation framework designed to assess U-MLLMs systematically. Our benchmark includes:
- Standardized Traditional Task Evaluation We sample from 12 datasets, covering 10 tasks with 30 subtasks, ensuring consistent and fair comparisons across studies.
- Unified Task Assessment We introduce five novel tasks testing multimodal reasoning, including image editing, commonsense QA with image generation, and geometric reasoning.
- Comprehensive Model Benchmarking We evaluate 12 leading U-MLLMs, such as Janus-Pro, EMU3, and VILA-U, alongside specialized understanding (e.g., Claude-3.5) and generation models (e.g., DALL-E-3).
Our findings reveal substantial performance gaps in existing U-MLLMs, highlighting the need for more robust models capable of handling mixed-modality tasks effectively.
