datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mllm-as-embodied-world-judge
MLLM-as-Embodied-World-Judge
Data for judging physical adherence and instruction alignment of generated
embodied-manipulation videos.
Start here
path
what it is
final/
the current release — train.jsonl (11,520), test.jsonl (802), and its README
data/
source and generated videos, referenced by video_url in the splits
Benchmark tooling
path
what it is
bench/LEADERBOARD.md
judge results table
bench/TESTSET.md
benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.MLLM_FeatMLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.MobileViews
🚀 MobileViews: A Large-Scale Mobile GUI Dataset
MobileViews is a large-scale dataset designed to support research on mobile agents and mobile user interface (UI) analysis. The first release, MobileViews-600K, includes over 600,000 mobile UI screenshot-view hierarchy (VH) pairs collected from over 20,000 apps on the Google Play Store. This dataset is based on the DroidBot, which we have optimized for large-scale data collection, capturing more comprehensive interaction details while… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/MobileViews.CL-VISTA
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
✨Introduction •
🥇 Methods Provided •
🏦 Benchmarks •
🎨 Models
🏃 How to run •
🤝 Acknowledgments •
🙂 Contact
If you like our project, please give us a star ⭐ on GitHub for the latest updates.
✨ Introduction
MCITlib is a unified library for continual instruction tuning of multimodal large language models (MLLMs). It integrates diverse continual learning methods into a… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/CL-VISTA.UCITUnofficial training-ready fork of HaiyangGuo/UCIT
DroidCall
DroidCall: A Dataset for LLM-powered Android Intent Invocation
paper|github
DroidCall is the first open-sourced, high-quality dataset designed for fine-tuning LLMs for accurate intent invocation on Android devices.
This repo contains data generated by DroidCall. The process of data generation is shown in the figure below
Details can be found in our paper and github repository.
What is Android Intent Invocation?
Android Intent is a key machanism in Android that allows… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/DroidCall.R1-Onevision
R1-Onevision
[📂 GitHub][📝 Paper]
[🤗 Reasoning Benchmark] [🤗 HF Demo]
R1-Onevision Dataset
Dataset Overview
The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities. Aimed at bridging the gap between visual and textual understanding, this dataset provides rich, context-aware reasoning tasks across diverse domains, including natural scenes, science, mathematical problems… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision.MLLMU-Bench
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
Abstract
Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.Domain40kMM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.MLLM-as-a-JudgeLogics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.MLLM-R1-Temp0227mllm-evalmllm-rlhf-testingMLLM-R1VTCTrainUse datasets>=4.0.0 to run the prepare code
EarthScience-MLLM-20K
EarthScience-MLLM-20K
A unified JSONL package for multimodal large-model training across three Earth-science domains:
Meteorology from ZhanxiangHua/WeatherQA_SFT.
Geography / map QA from HuggingFaceM4/the_cauldron config mapqa.
Remote-sensing common-sense QA + grounding/detection from xiang709/VRSBench.
The package intentionally excludes segmentation-style targets. Each JSONL line is one training/evaluation unit.
Files
train.jsonl: 20000 examples.
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-MLLM-20K.DCLVTCBench
Dataset Card for VTCBench
Vision-Text Compression Benchmark (VTCBench)
revisits Needle-In-A-Haystack (NIAH)
from a VLM's perspective by converting long context into rendered images.
This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and
memorize long context as images. Specifically, this benchmark includes 3 tasks:
Retrieval: Vision-NIAH VQA task for information retrieval and aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/VTCBench.sarab
Sarab Dataset
The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination
evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD
benchmark. Code and evaluation scripts are on
GitHub.
What this is
A human-captioned pool of Arabic Cultural Visual Vocabulary (ACVV) images
(architecture, attire, cuisine, objects, script), built into five evaluation
modes:
base — plain image, direct Arabic question.
sec (specious context) — image… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.MLLM_pathRefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.ReasonMed
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
📄 Paper |
💻 Code |
📊 Dataset
ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.AudioQA-1MMLLM-CL
MLLM-CL: Continual Learning for Multimodal Large Language Models
This is the official dataset repository for MLLM-CL: Continual Learning for Multimodal Large Language Models.
Paper: MLLM-CL: Continual Learning for Multimodal Large Language Models
Code: https://github.com/bjzhb666/MLLM-CL
Recent Multimodal Large Language Models (MLLMs) excel in vision-language understanding but face challenges in adapting to dynamic real-world scenarios that require continuous integration of new… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/MLLM-CL.Logics-STEM-SFT-Dataset-Open-5.3MOmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.
