datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.multimodal-time-series-forecastingmultimodal-grounding-ooc
Multimodal Grounding of Explanations for Out-of-Context Misinformation Detection
This dataset contains the outputs, explanations, and visual grounding audits for three vision-language model configurations evaluated on out-of-context (OOC) misinformation detection:
Gemma-4-31B-It (Direct): Baseline API evaluation with minimal thinking compute.
Gemma-4-31B-It (Thinking): Deliberation API evaluation with high thinking compute (up to 4,352 tokens).
Gemma-3-27B-It (Direct):… See the full description on the dataset page: https://huggingface.co/datasets/jordansp/multimodal-grounding-ooc.NLS-CH-Multimodal
NLS-CH-Multimodal: A Large-Scale Multi-Modal Cultural Heritage Corpus
Dataset Summary
NLS-CH-Multimodal is a large-scale multimodal corpus derived from the National Library of Scotland (NLS) digital collections,
comprising over 512,000 files and exceeding 1 TB in size. The dataset is designed to support Information Retrieval (IR),
Retrieval-Augmented Generation (RAG), and analysis of large language model (LLM) behaviour on historical data.
The corpus integrates… See the full description on the dataset page: https://huggingface.co/datasets/NeuraSearchLab/NLS-CH-Multimodal.mirror
MIRROR Dataset
MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance.
Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
The dataset includes:
Client profile metadata (CACTUS idx, CelebA idx)
Dialogue written in a screenplay format, including stage directions that describe facial expressions
⚠️ Images themselves are not included to comply with the CelebA license.
However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.hotel-multimodalmedical-pathology-radiology-multimodal
🚀 Datavendor Multimodal Medical Dataset (Radiology & Pathology)
A premium, expert-validated, multi-modal clinical dataset engineered for high-accuracy AI and machine learning applications. This repository provides structured sample data for evaluation and integration testing.
1. Structured Patient Context (SPC) – Radiology Reports
We have engineered a large-scale cohort of radiology reports into a Structured Patient Context (SPC) format, enabling efficient downstream AI… See the full description on the dataset page: https://huggingface.co/datasets/datavendor/medical-pathology-radiology-multimodal.amv_genre_multimodal_dataset
Estrutura do Conjunto de Dados
O conjunto de dados, em formato CSV, foi preparado usando técnicas de visualização de mídia (media visualization). Ele é composto por vídeos de humor sobre temas sociopolíticos, coletados do YouTube. O arquivo contém diversas colunas para uma análise multimodal:
amv_genre: O gênero do vídeo, categorizado como drama ou action.
Características Visuais: Métricas de cor e brilho, como a média, mediana, desvio padrão e frequência dominante para matiz… See the full description on the dataset page: https://huggingface.co/datasets/Dumoura/amv_genre_multimodal_dataset.Multimodal-InfomaxMulti-Modal_Sentiment_Analysis_in_E-commerce
