datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.hle-multimodalmultimodal-LLMs-See-Sentiment
MLLMsent — datasets and experiment results
Every input and every output of "Multimodal LLMs See Sentiment"
(arXiv:2508.16873): the image descriptions generated by six multimodal
LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete
per-fold results of all 141 experiments.
Paper: arXiv:2508.16873
Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.panda-70m
Panda 70M dataset by Snap Inc
70M video-caption pairs
Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading
IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.multimodal-time-series-forecastingmultimodal_query_rewrites
ReVision: Visual Instruction Rewriting Dataset
Dataset Summary
The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises:
Image: A visual scene containing relevant information.
Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/anonymoususerrevision/multimodal_query_rewrites.multimodal_query_rewrites
ReVision: Visual Instruction Rewriting Dataset
Dataset Summary
The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises:
Image: A visual scene containing relevant information.
Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/hsiangfu/multimodal_query_rewrites.multimodal-grounding-ooc
Multimodal Grounding of Explanations for Out-of-Context Misinformation Detection
This dataset contains the outputs, explanations, and visual grounding audits for three vision-language model configurations evaluated on out-of-context (OOC) misinformation detection:
Gemma-4-31B-It (Direct): Baseline API evaluation with minimal thinking compute.
Gemma-4-31B-It (Thinking): Deliberation API evaluation with high thinking compute (up to 4,352 tokens).
Gemma-3-27B-It (Direct):… See the full description on the dataset page: https://huggingface.co/datasets/jordansp/multimodal-grounding-ooc.NLS-CH-Multimodal
NLS-CH-Multimodal: A Large-Scale Multi-Modal Cultural Heritage Corpus
Dataset Summary
NLS-CH-Multimodal is a large-scale multimodal corpus derived from the National Library of Scotland (NLS) digital collections,
comprising over 512,000 files and exceeding 1 TB in size. The dataset is designed to support Information Retrieval (IR),
Retrieval-Augmented Generation (RAG), and analysis of large language model (LLM) behaviour on historical data.
The corpus integrates… See the full description on the dataset page: https://huggingface.co/datasets/NeuraSearchLab/NLS-CH-Multimodal.Urdu-Multimodal-Emotion-Datasetmultimodal-constraint-evals
Multimodal Constraint Evaluation Dataset
This dataset contains human-designed evaluation cases for multimodal image generation models.
Purpose
The goal of this dataset is to expose repeatable failure modes related to:
object counting under strict constraints,
loss of uniqueness across generated entities,
layout and panel consistency,
multi-step and multi-surface reasoning,
planning vs rendering behavior in single-pass generation.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/multimodal-constraint-evals.test-audio-datasetgsl-multimodal-annotation
GSL Multimodal Annotation Dataset
Multimodal annotation of 20 signs from Ghanaian Sign Language (GSL),
capturing manual and non-manual phonological features across 14 columns
including handshape, location, movement, facial expression, mouth
pattern, and head movement.
Dataset description
This dataset accompanies a pilot Linked Data representation of GSL
(DOI: 10.5281/zenodo.20961293). It documents the annotation decisions,
uncertainties, and limitations… See the full description on the dataset page: https://huggingface.co/datasets/LINGUISTEUNICE/gsl-multimodal-annotation.mirror
MIRROR Dataset
MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance.
Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
The dataset includes:
Client profile metadata (CACTUS idx, CelebA idx)
Dialogue written in a screenplay format, including stage directions that describe facial expressions
⚠️ Images themselves are not included to comply with the CelebA license.
However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.hotel-multimodalmultimodalLPT2medical-pathology-radiology-multimodal
🚀 Datavendor Multimodal Medical Dataset (Radiology & Pathology)
A premium, expert-validated, multi-modal clinical dataset engineered for high-accuracy AI and machine learning applications. This repository provides structured sample data for evaluation and integration testing.
1. Structured Patient Context (SPC) – Radiology Reports
We have engineered a large-scale cohort of radiology reports into a Structured Patient Context (SPC) format, enabling efficient downstream AI… See the full description on the dataset page: https://huggingface.co/datasets/datavendor/medical-pathology-radiology-multimodal.multimodal-emotion-dataset
Multimodal Emotion Recognition Dataset
This dataset contains 1000 samples of short text messages with associated emojis and sentiment labels. It is designed for fine-tuning models on multimodal emotion recognition tasks, where both text and emojis contribute to predicting the emotion.
Dataset Fields
text: A short sentence or phrase.
emojis: A list of emojis used in the text.
sentiment: Sentiment label with three possible values:
positive
negative
neutral… See the full description on the dataset page: https://huggingface.co/datasets/prajubhao/multimodal-emotion-dataset.multimodal_human_computer_interaction_gesturesRoomly-Student-Bios-Multimodal
Roomly: Multimodal Roommate Matching Dataset
🎯 Problem Statement
Finding a roommate is often reduced to dry filters like "budget" and "location". Roomly aims to revolutionize this by focusing on personality, lifestyle, and visual preferences. This dataset provides synthetic student profiles and their ideal room environments.
📊 Exploratory Data Analysis (EDA)
1. User Persona Distribution
Our dataset contains a balanced mix of different student… See the full description on the dataset page: https://huggingface.co/datasets/Orib24/Roomly-Student-Bios-Multimodal.amv_genre_multimodal_dataset
Estrutura do Conjunto de Dados
O conjunto de dados, em formato CSV, foi preparado usando técnicas de visualização de mídia (media visualization). Ele é composto por vídeos de humor sobre temas sociopolíticos, coletados do YouTube. O arquivo contém diversas colunas para uma análise multimodal:
amv_genre: O gênero do vídeo, categorizado como drama ou action.
Características Visuais: Métricas de cor e brilho, como a média, mediana, desvio padrão e frequência dominante para matiz… See the full description on the dataset page: https://huggingface.co/datasets/Dumoura/amv_genre_multimodal_dataset.MultiModal-Planning-BenchmarkMulti-Modal_Sentiment_Analysis_in_E-commercehumanoid-multimodal-intent-resolution
Multimodal Intent Resolution Dataset
Maps multi-sensor inputs into unified humanoid intents.
Multimodal-Infomaximage_caption_pairs_for_multimodal
