MLLM
Datasets
All datasets matching “MLLM”mllm-as-embodied-world-judge
MLLM-as-Embodied-World-Judge
Data for judging physical adherence and instruction alignment of generated
embodied-manipulation videos.
Start here
path
what it is
final/
the current release — train.jsonl (11,520), test.jsonl (802), and its README
data/
source and generated videos, referenced by video_url in the splits
Benchmark tooling
path
what it is
bench/LEADERBOARD.md
judge results table
bench/TESTSET.md
benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.MLLM_FeatMLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.MobileViews
🚀 MobileViews: A Large-Scale Mobile GUI Dataset
MobileViews is a large-scale dataset designed to support research on mobile agents and mobile user interface (UI) analysis. The first release, MobileViews-600K, includes over 600,000 mobile UI screenshot-view hierarchy (VH) pairs collected from over 20,000 apps on the Google Play Store. This dataset is based on the DroidBot, which we have optimized for large-scale data collection, capturing more comprehensive interaction details while… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/MobileViews.CL-VISTA
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
✨Introduction •
🥇 Methods Provided •
🏦 Benchmarks •
🎨 Models
🏃 How to run •
🤝 Acknowledgments •
🙂 Contact
If you like our project, please give us a star ⭐ on GitHub for the latest updates.
✨ Introduction
MCITlib is a unified library for continual instruction tuning of multimodal large language models (MLLMs). It integrates diverse continual learning methods into a… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/CL-VISTA.
