datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.CL-VISTA
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
✨Introduction •
🥇 Methods Provided •
🏦 Benchmarks •
🎨 Models
🏃 How to run •
🤝 Acknowledgments •
🙂 Contact
If you like our project, please give us a star ⭐ on GitHub for the latest updates.
✨ Introduction
MCITlib is a unified library for continual instruction tuning of multimodal large language models (MLLMs). It integrates diverse continual learning methods into a… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/CL-VISTA.UCITUnofficial training-ready fork of HaiyangGuo/UCIT
DroidCall
DroidCall: A Dataset for LLM-powered Android Intent Invocation
paper|github
DroidCall is the first open-sourced, high-quality dataset designed for fine-tuning LLMs for accurate intent invocation on Android devices.
This repo contains data generated by DroidCall. The process of data generation is shown in the figure below
Details can be found in our paper and github repository.
What is Android Intent Invocation?
Android Intent is a key machanism in Android that allows… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/DroidCall.R1-Onevision
R1-Onevision
[📂 GitHub][📝 Paper]
[🤗 Reasoning Benchmark] [🤗 HF Demo]
R1-Onevision Dataset
Dataset Overview
The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities. Aimed at bridging the gap between visual and textual understanding, this dataset provides rich, context-aware reasoning tasks across diverse domains, including natural scenes, science, mathematical problems… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision.MLLMU-Bench
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
Abstract
Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.Domain40kLogics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.mllm-evalDCLVTCBench
Dataset Card for VTCBench
Vision-Text Compression Benchmark (VTCBench)
revisits Needle-In-A-Haystack (NIAH)
from a VLM's perspective by converting long context into rendered images.
This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and
memorize long context as images. Specifically, this benchmark includes 3 tasks:
Retrieval: Vision-NIAH VQA task for information retrieval and aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/VTCBench.sarab
Sarab Dataset
The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination
evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD
benchmark. Code and evaluation scripts are on
GitHub.
What this is
A human-captioned pool of Arabic Cultural Visual Vocabulary (ACVV) images
(architecture, attire, cuisine, objects, script), built into five evaluation
modes:
base — plain image, direct Arabic question.
sec (specious context) — image… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.VTCTrainUse datasets>=4.0.0 to run the prepare code
RefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.ReasonMed
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
📄 Paper |
💻 Code |
📊 Dataset
ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.Logics-STEM-SFT-Dataset-Open-5.3MOmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.mllm-shap
MLLM-SHAP experiment datasets
Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora.
Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench).
Quick load
Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.DCL-10PercentComic-9K
Comic-9K
Image
Extracting all images.
cat images.tar.gz.aa images.tar.gz.ab images.tar.gz.ac images.tar.gz.ad images.tar.gz.ae > images.tar.gz
tar xvzf images.tar.gz
Summary
We provide human-written plot synopsis.
summary.jsonl
COCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/COCO.ReferringImageCaptioningPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/ReferringImageCaptioning.MLLMGuardmllm_cl_textvqaMLLM-CL-ReplayData
MLLM-CL: Continual Learning for Multimodal Large Language Models
This is the official dataset repository of MLLM-CL and MR-LoRA. MLLM-CL is a novel benchmark encompassing domain and ability continual learning, where the former focuses on independently and identically distributed (IID) evaluation across evolving mainstream domains, whereas the latter evaluates on non-IID scenarios with emerging model ability. MR-LoRA prevents catastrophic interference through parameter isolation and… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/MLLM-CL-ReplayData.R1-Onevision-Bench
R1-Onevision-Bench
[📂 GitHub][📝 Paper]
[🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo]
Dataset Overview
R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.mllm_cl_vizwizLogics-SWE-Env-2.5K
Logics-SWE-Env-2.5K
2,553 software engineering task instances · 1,771 repositories · 4 programming languages
🤗 Related model: Logics-SWE-Qwen3.6-27B
📄 Paper: One to More, More to One
Overview
What is this dataset?
Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771 GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.Long-VITA-Data
