datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLLM-as-a-JudgeMindCube
MindCube: Spatial Mental Modeling from Limited Views
MindCube is a novel benchmark designed to evaluate how well Vision Language Models (VLMs) can form robust spatial mental models from limited visual views. It comprises 21,154 questions across 3,268 images, assessing capabilities such as cognitive mapping (representing positions), perspective-taking (orientations), and mental simulation (dynamics for "what-if" movements). The dataset aims to expose critical gaps in existing VLMs'… See the full description on the dataset page: https://huggingface.co/datasets/MLL-Lab/MindCube.MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.MultiBBQ
MultiBBQ: A Fairness Benchmark for Multimodal LLMs
Controllable diagnosis of social bias in multimodal LLMs with synthetic images.
MultiBBQ is a fairness evaluation benchmark for multimodal large language models (MLLMs).
It extends the language-only BBQ benchmark into the visual
domain: each attested social bias is paired with an AI-generated photorealistic image of two
people who differ only in the target demographic, so a model's fairness can be… See the full description on the dataset page: https://huggingface.co/datasets/MLL-Lab/MultiBBQ.MLLM-Fabric
🧵 MLLM-Fabric: Multimodal LLM-Driven Robotic Framework for Fabric Sorting and Selection
📄 Overview
This is the official repository for the paper:
MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
Accepted to IEEE Robotics and Automation Letters (RA-L)
🏫 About This Work
This work is from the Robot-Assisted Living LAboratory (RALLA) at the University of York, UK.
🧵 Fabric Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/EuniceF/MLLM-Fabric.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096
Derived dataset note
This dataset was derived from OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking as a part of arxiv.org/abs/2603.22276.
Field changes:
question -> query
qwen3vl_235b_thinking_response -> response
image -> images (single-item list)
added tok_len, computed with tokenizer Qwen/Qwen3-8B on query + '\n\n' + response
add_special_tokens=False
The original README content is preserved below.
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/eyes-ml/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096.GeneralScience-MLLM-22K
GeneralScience-MLLM-22K
Dataset Summary
GeneralScience-MLLM-22K is a unified general-science multiple-choice QA collection built from local snapshots of SciQ, AI2 ARC, and ScienceQA. It follows the same release style as a subject-specific MLLM dataset: every sample is stored as one JSONL record, text-only and image-text examples share one schema, and ScienceQA images are exported as standalone files referenced by relative paths.
The release contains 22,661… See the full description on the dataset page: https://huggingface.co/datasets/gineven/GeneralScience-MLLM-22K.mlp
1. Overview of ViRL39K
ViRL39K (pronounced as "viral") provides a curated collection of 38,870 verifiable QAs for Vision-Language RL training.
It is built on top of newly collected problems and existing datasets (
Llava-OneVision,
R1-OneVision,
MM-Eureka,
MM-Math,
M3CoT,
DeepScaleR,
MV-Math)
through cleaning, reformatting, rephrasing and verification.ViRL39K lays the foundation for SoTA Vision-Language Reasoning Model VL-Rethinker. It has the following merits:
high-quality and… See the full description on the dataset page: https://huggingface.co/datasets/vasuverma/mlp.Amazon_ml_challenge_flitered_dataset
