CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFriends /mllm-as-embodied-world-judge MLLM-as-Embodied-World-Judge Data for judging physical adherence and instruction alignment of generated embodied-manipulation videos. Start here path what it is final/ the current release — train.jsonl (11,520), test.jsonl (802), and its README data/ source and generated videos, referenced by video_url in the splits Benchmark tooling path what it is bench/LEADERBOARD.md judge results table bench/TESTSET.md benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.2 likes17k downloads15d agoHugging Face02lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B7 likes2.9k downloads22d agoHugging Face03ElliotJenkins /MLLM_Feat0 likes2.5k downloads1y agoHugging Face04zr-zhang /MLLM-Generated-Image-Detection-Dataset MLLM-Generated Image Dataset This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection. Paper | Code Dataset Summary We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.imageimage-classification1K<n<10K1 likes2.1k downloads2mo agoHugging Face05mllmTeam /MobileViews 🚀 MobileViews: A Large-Scale Mobile GUI Dataset MobileViews is a large-scale dataset designed to support research on mobile agents and mobile user interface (UI) analysis. The first release, MobileViews-600K, includes over 600,000 mobile UI screenshot-view hierarchy (VH) pairs collected from over 20,000 apps on the Google Play Store. This dataset is based on the DroidBot, which we have optimized for large-scale data collection, capturing more comprehensive interaction details while… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/MobileViews.question-answering45 likes2.1k downloads2y agoHugging Face06MLLM-CL /CL-VISTA MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark ✨Introduction • 🥇 Methods Provided • 🏦 Benchmarks • 🎨 Models 🏃 How to run • 🤝 Acknowledgments • 🙂 Contact If you like our project, please give us a star ⭐ on GitHub for the latest updates. ✨ Introduction MCITlib is a unified library for continual instruction tuning of multimodal large language models (MLLMs). It integrates diverse continual learning methods into a… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/CL-VISTA.text100K<n<1M1 likes1.7k downloads4mo agoHugging Face07MLLM-CL /UCITUnofficial training-ready fork of HaiyangGuo/UCIT image100K<n<1M1 likes1.4k downloads6mo agoHugging Face08mllmTeam /DroidCall DroidCall: A Dataset for LLM-powered Android Intent Invocation paper|github DroidCall is the first open-sourced, high-quality dataset designed for fine-tuning LLMs for accurate intent invocation on Android devices. This repo contains data generated by DroidCall. The process of data generation is shown in the figure below Details can be found in our paper and github repository. What is Android Intent Invocation? Android Intent is a key machanism in Android that allows… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/DroidCall.texttext-generation10K<n<100K3 likes1.1k downloads2y agoHugging Face09Fancy-MLLM /R1-Onevision R1-Onevision [📂 GitHub][📝 Paper] [🤗 Reasoning Benchmark] [🤗 HF Demo] R1-Onevision Dataset Dataset Overview The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities. Aimed at bridging the gap between visual and textual understanding, this dataset provides rich, context-aware reasoning tasks across diverse domains, including natural scenes, science, mathematical problems… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision.textquestion-answering100K<n<1M47 likes949 downloads2y agoHugging Face10MLLMMU /MLLMU-Bench Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench Abstract Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.image1K<n<10K6 likes925 downloads2y agoHugging Face11MLLM-CL /Domain40kimage100K<n<1M1 likes870 downloads6mo agoHugging Face12EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes825 downloads6mo agoHugging Face13ONE-Lab /MLLM-as-a-Judgeimagequestion-answering1K<n<10K4 likes821 downloads2y agoHugging Face14Logics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes809 downloads8mo agoHugging Face15Wendy-Fly /MLLM-R1-Temp02270 likes772 downloads1y agoHugging Face16qinbright99 /mllm-evaltext100K<n<1M0 likes739 downloads1y agoHugging Face17ZoeyZou /mllm-rlhf-testing0 likes731 downloads12d agoHugging Face18Wendy-Fly /MLLM-R10 likes611 downloads1y agoHugging Face19MLLM-CL /VTCTrainUse datasets>=4.0.0 to run the prepare code text10K<n<100K0 likes577 downloads7mo agoHugging Face20moTcream /EarthScience-MLLM-20K EarthScience-MLLM-20K A unified JSONL package for multimodal large-model training across three Earth-science domains: Meteorology from ZhanxiangHua/WeatherQA_SFT. Geography / map QA from HuggingFaceM4/the_cauldron config mapqa. Remote-sensing common-sense QA + grounding/detection from xiang709/VRSBench. The package intentionally excludes segmentation-style targets. Each JSONL line is one training/evaluation unit. Files train.jsonl: 20000 examples. test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-MLLM-20K.imagevisual-question-answering10K<n<100K0 likes571 downloads2mo agoHugging Face21MLLM-CL /DCLimage100K<n<1M0 likes480 downloads4mo agoHugging Face22MLLM-CL /VTCBench Dataset Card for VTCBench Vision-Text Compression Benchmark (VTCBench) revisits Needle-In-A-Haystack (NIAH) from a VLM's perspective by converting long context into rendered images. This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and memorize long context as images. Specifically, this benchmark includes 3 tasks: Retrieval: Vision-NIAH VQA task for information retrieval and aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/VTCBench.imagevisual-question-answering1K<n<10K4 likes448 downloads28d agoHugging Face23Sarab-MLLMs /sarab Sarab Dataset The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD benchmark. Code and evaluation scripts are on GitHub. What this is A human-captioned pool of Arabic Cultural Visual Vocabulary (ACVV) images (architecture, attire, cuisine, objects, script), built into five evaluation modes: base — plain image, direct Arabic question. sec (specious context) — image… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.imagevisual-question-answeringn<1K0 likes427 downloads13d agoHugging Face24ElliotJenkins /MLLM_pathtabularn<1K0 likes375 downloads1y agoHugging Face25PaDT-MLLM /RefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.textobject-detection100K<n<1M4 likes365 downloads1y agoHugging Face26lingshu-medical-mllm /ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 📄 Paper  |  💻 Code  |  📊 Dataset ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.textquestion-answering1M<n<10M95 likes347 downloads1y agoHugging Face27VITA-MLLM /AudioQA-1M3 likes338 downloads2y agoHugging Face28MLLM-CL /MLLM-CL MLLM-CL: Continual Learning for Multimodal Large Language Models This is the official dataset repository for MLLM-CL: Continual Learning for Multimodal Large Language Models. Paper: MLLM-CL: Continual Learning for Multimodal Large Language Models Code: https://github.com/bjzhb666/MLLM-CL Recent Multimodal Large Language Models (MLLMs) excel in vision-language understanding but face challenges in adapting to dynamic real-world scenarios that require continuous integration of new… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/MLLM-CL.image-text-to-text100K<n<1M2 likes313 downloads4mo agoHugging Face29Logics-MLLM /Logics-STEM-SFT-Dataset-Open-5.3Mtext1M<n<10M4 likes279 downloads8mo agoHugging Face30Logics-MLLM /OmniParsingBench 🤗 Model   |   📑 Technical Report   |   💻 GitHub OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities. Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.image1K<n<10K2 likes271 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.