CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.1k downloads22d agoHugging Face02MLLM-CL /CL-VISTA MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark ✨Introduction • 🥇 Methods Provided • 🏦 Benchmarks • 🎨 Models 🏃 How to run • 🤝 Acknowledgments • 🙂 Contact If you like our project, please give us a star ⭐ on GitHub for the latest updates. ✨ Introduction MCITlib is a unified library for continual instruction tuning of multimodal large language models (MLLMs). It integrates diverse continual learning methods into a… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/CL-VISTA.text100K<n<1M1 likes1.7k downloads4mo agoHugging Face03MLLM-CL /UCITUnofficial training-ready fork of HaiyangGuo/UCIT image100K<n<1M1 likes1.2k downloads6mo agoHugging Face04mllmTeam /DroidCall DroidCall: A Dataset for LLM-powered Android Intent Invocation paper|github DroidCall is the first open-sourced, high-quality dataset designed for fine-tuning LLMs for accurate intent invocation on Android devices. This repo contains data generated by DroidCall. The process of data generation is shown in the figure below Details can be found in our paper and github repository. What is Android Intent Invocation? Android Intent is a key machanism in Android that allows… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/DroidCall.texttext-generation10K<n<100K3 likes1.1k downloads2y agoHugging Face05Fancy-MLLM /R1-Onevision R1-Onevision [📂 GitHub][📝 Paper] [🤗 Reasoning Benchmark] [🤗 HF Demo] R1-Onevision Dataset Dataset Overview The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities. Aimed at bridging the gap between visual and textual understanding, this dataset provides rich, context-aware reasoning tasks across diverse domains, including natural scenes, science, mathematical problems… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision.textquestion-answering100K<n<1M47 likes968 downloads2y agoHugging Face06MLLMMU /MLLMU-Bench Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench Abstract Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.image1K<n<10K6 likes954 downloads2y agoHugging Face07MLLM-CL /Domain40kimage100K<n<1M1 likes869 downloads6mo agoHugging Face08Logics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes829 downloads8mo agoHugging Face09EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes802 downloads6mo agoHugging Face10qinbright99 /mllm-evaltext100K<n<1M0 likes739 downloads1y agoHugging Face11MLLM-CL /DCLimage100K<n<1M0 likes480 downloads4mo agoHugging Face12MLLM-CL /VTCBench Dataset Card for VTCBench Vision-Text Compression Benchmark (VTCBench) revisits Needle-In-A-Haystack (NIAH) from a VLM's perspective by converting long context into rendered images. This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and memorize long context as images. Specifically, this benchmark includes 3 tasks: Retrieval: Vision-NIAH VQA task for information retrieval and aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/VTCBench.imagevisual-question-answering1K<n<10K4 likes452 downloads29d agoHugging Face13Sarab-MLLMs /sarab Sarab Dataset The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD benchmark. Code and evaluation scripts are on GitHub. What this is A human-captioned pool of Arabic Cultural Visual Vocabulary (ACVV) images (architecture, attire, cuisine, objects, script), built into five evaluation modes: base — plain image, direct Arabic question. sec (specious context) — image… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.imagevisual-question-answeringn<1K0 likes430 downloads14d agoHugging Face14MLLM-CL /VTCTrainUse datasets>=4.0.0 to run the prepare code text10K<n<100K0 likes399 downloads7mo agoHugging Face15PaDT-MLLM /RefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.textobject-detection100K<n<1M4 likes357 downloads1y agoHugging Face16lingshu-medical-mllm /ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 📄 Paper  |  💻 Code  |  📊 Dataset ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.textquestion-answering1M<n<10M95 likes344 downloads1y agoHugging Face17Logics-MLLM /Logics-STEM-SFT-Dataset-Open-5.3Mtext1M<n<10M4 likes278 downloads8mo agoHugging Face18Logics-MLLM /OmniParsingBench 🤗 Model   |   📑 Technical Report   |   💻 GitHub OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities. Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.image1K<n<10K2 likes277 downloads6mo agoHugging Face19Pawlo77 /mllm-shap MLLM-SHAP experiment datasets Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora. Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench). Quick load Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.tabulartext-generation1K<n<10K2 likes197 downloads4mo agoHugging Face20MLLM-CL /DCL-10Percentimage10K<n<100K0 likes166 downloads6mo agoHugging Face21VITA-MLLM /Comic-9K Comic-9K Image Extracting all images. cat images.tar.gz.aa images.tar.gz.ab images.tar.gz.ac images.tar.gz.ad images.tar.gz.ae > images.tar.gz tar xvzf images.tar.gz Summary We provide human-written plot synopsis. summary.jsonl image100K<n<1M6 likes158 downloads2y agoHugging Face22PaDT-MLLM /COCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/COCO.textobject-detection100K<n<1M1 likes147 downloads1y agoHugging Face23PaDT-MLLM /ReferringImageCaptioningPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/ReferringImageCaptioning.textimage-to-text100K<n<1M3 likes138 downloads1y agoHugging Face24Yujie-AI /MLLMGuardimage1K<n<10K0 likes131 downloads22d agoHugging Face25zhhxte /mllm_cl_textvqaimage10K<n<100K0 likes120 downloads1y agoHugging Face26MLLM-CL /MLLM-CL-ReplayData MLLM-CL: Continual Learning for Multimodal Large Language Models This is the official dataset repository of MLLM-CL and MR-LoRA. MLLM-CL is a novel benchmark encompassing domain and ability continual learning, where the former focuses on independently and identically distributed (IID) evaluation across evolving mainstream domains, whereas the latter evaluates on non-IID scenarios with emerging model ability. MR-LoRA prevents catastrophic interference through parameter isolation and… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/MLLM-CL-ReplayData.textimage-text-to-text100K<n<1M0 likes120 downloads1y agoHugging Face27Fancy-MLLM /R1-Onevision-Bench R1-Onevision-Bench [📂 GitHub][📝 Paper] [🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo] Dataset Overview R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.textquestion-answeringn<1K3 likes117 downloads2y agoHugging Face28zhhxte /mllm_cl_vizwizimage10K<n<100K0 likes116 downloads1y agoHugging Face29Logics-MLLM /Logics-SWE-Env-2.5K Logics-SWE-Env-2.5K 2,553 software engineering task instances · 1,771 repositories · 4 programming languages 🤗 Related model: Logics-SWE-Qwen3.6-27B 📄 Paper: One to More, More to One Overview What is this dataset? Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771 GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.tabular1K<n<10K2 likes103 downloads18h agoHugging Face30VITA-MLLM /Long-VITA-Datatext10M<n<100M2 likes96 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.