CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K3 likes8.4k downloads5mo agoHugging Face02OptimusePrime /hle-multimodaltextn<1K0 likes1.4k downloads1y agoHugging Face03neemiasbsilva /multimodal-LLMs-See-Sentiment MLLMsent — datasets and experiment results Every input and every output of "Multimodal LLMs See Sentiment" (arXiv:2508.16873): the image descriptions generated by six multimodal LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete per-fold results of all 141 experiments. Paper: arXiv:2508.16873 Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.texttext-classification10K<n<100K1 likes359 downloads1mo agoHugging Face04trucyberlab /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M0 likes233 downloads3mo agoHugging Face05multimodalart /panda-70m Panda 70M dataset by Snap Inc 70M video-caption pairs Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading textimage-to-text100K<n<1M15 likes223 downloads2y agoHugging Face06alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes199 downloads4mo agoHugging Face07harryCJ /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M0 likes64 downloads2mo agoHugging Face08FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes61 downloads5mo agoHugging Face09illilliks /multimodal-time-series-forecastingimage100K<n<1M1 likes59 downloads1y agoHugging Face10anonymoususerrevision /multimodal_query_rewrites ReVision: Visual Instruction Rewriting Dataset Dataset Summary The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises: Image: A visual scene containing relevant information. Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/anonymoususerrevision/multimodal_query_rewrites.image10K<n<100K1 likes46 downloads2y agoHugging Face11hsiangfu /multimodal_query_rewrites ReVision: Visual Instruction Rewriting Dataset Dataset Summary The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises: Image: A visual scene containing relevant information. Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/hsiangfu/multimodal_query_rewrites.image10K<n<100K0 likes37 downloads2y agoHugging Face12jordansp /multimodal-grounding-ooc Multimodal Grounding of Explanations for Out-of-Context Misinformation Detection This dataset contains the outputs, explanations, and visual grounding audits for three vision-language model configurations evaluated on out-of-context (OOC) misinformation detection: Gemma-4-31B-It (Direct): Baseline API evaluation with minimal thinking compute. Gemma-4-31B-It (Thinking): Deliberation API evaluation with high thinking compute (up to 4,352 tokens). Gemma-3-27B-It (Direct):… See the full description on the dataset page: https://huggingface.co/datasets/jordansp/multimodal-grounding-ooc.tabularimage-to-text1K<n<10K0 likes30 downloads2mo agoHugging Face13NeuraSearchLab /NLS-CH-Multimodal NLS-CH-Multimodal: A Large-Scale Multi-Modal Cultural Heritage Corpus Dataset Summary NLS-CH-Multimodal is a large-scale multimodal corpus derived from the National Library of Scotland (NLS) digital collections, comprising over 512,000 files and exceeding 1 TB in size. The dataset is designed to support Information Retrieval (IR), Retrieval-Augmented Generation (RAG), and analysis of large language model (LLM) behaviour on historical data. The corpus integrates… See the full description on the dataset page: https://huggingface.co/datasets/NeuraSearchLab/NLS-CH-Multimodal.tabularn<1K0 likes27 downloads4mo agoHugging Face14Maisum-Abbas-123 /Urdu-Multimodal-Emotion-Datasetaudio1K<n<10K0 likes25 downloads8mo agoHugging Face158Planetterraforming /multimodal-constraint-evals Multimodal Constraint Evaluation Dataset This dataset contains human-designed evaluation cases for multimodal image generation models. Purpose The goal of this dataset is to expose repeatable failure modes related to: object counting under strict constraints, loss of uniqueness across generated entities, layout and panel consistency, multi-step and multi-surface reasoning, planning vs rendering behavior in single-pass generation. Dataset Structure The dataset… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/multimodal-constraint-evals.texttext-to-imagen<1K0 likes20 downloads8mo agoHugging Face16multimodalart /test-audio-datasettextn<1K0 likes18 downloads2y agoHugging Face17LINGUISTEUNICE /gsl-multimodal-annotation GSL Multimodal Annotation Dataset Multimodal annotation of 20 signs from Ghanaian Sign Language (GSL), capturing manual and non-manual phonological features across 14 columns including handshape, location, movement, facial expression, mouth pattern, and head movement. Dataset description This dataset accompanies a pilot Linked Data representation of GSL (DOI: 10.5281/zenodo.20961293). It documents the annotation decisions, uncertainties, and limitations… See the full description on the dataset page: https://huggingface.co/datasets/LINGUISTEUNICE/gsl-multimodal-annotation.textn<1K0 likes15 downloads3mo agoHugging Face18multimodal-reframing /mirror MIRROR Dataset MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance. Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance The dataset includes: Client profile metadata (CACTUS idx, CelebA idx) Dialogue written in a screenplay format, including stage directions that describe facial expressions ⚠️ Images themselves are not included to comply with the CelebA license. However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.tabulartext-generationn<1K2 likes14 downloads10mo agoHugging Face19ruslanmv /hotel-multimodalimage1M<n<10M2 likes12 downloads2y agoHugging Face20machinev /multimodalLPT2textn<1K0 likes9 downloads2y agoHugging Face21datavendor /medical-pathology-radiology-multimodalgated 🚀 Datavendor Multimodal Medical Dataset (Radiology & Pathology) A premium, expert-validated, multi-modal clinical dataset engineered for high-accuracy AI and machine learning applications. This repository provides structured sample data for evaluation and integration testing. 1. Structured Patient Context (SPC) – Radiology Reports We have engineered a large-scale cohort of radiology reports into a Structured Patient Context (SPC) format, enabling efficient downstream AI… See the full description on the dataset page: https://huggingface.co/datasets/datavendor/medical-pathology-radiology-multimodal.tabulartext-classificationn<1K3 likes9 downloads6mo agoHugging Face22prajubhao /multimodal-emotion-dataset Multimodal Emotion Recognition Dataset This dataset contains 1000 samples of short text messages with associated emojis and sentiment labels. It is designed for fine-tuning models on multimodal emotion recognition tasks, where both text and emojis contribute to predicting the emotion. Dataset Fields text: A short sentence or phrase. emojis: A list of emojis used in the text. sentiment: Sentiment label with three possible values: positive negative neutral… See the full description on the dataset page: https://huggingface.co/datasets/prajubhao/multimodal-emotion-dataset.text1K<n<10K0 likes8 downloads2y agoHugging Face23Shoriful025 /multimodal_human_computer_interaction_gesturestextn<1K0 likes6 downloads9mo agoHugging Face24Orib24 /Roomly-Student-Bios-Multimodal Roomly: Multimodal Roommate Matching Dataset 🎯 Problem Statement Finding a roommate is often reduced to dry filters like "budget" and "location". Roomly aims to revolutionize this by focusing on personality, lifestyle, and visual preferences. This dataset provides synthetic student profiles and their ideal room environments. 📊 Exploratory Data Analysis (EDA) 1. User Persona Distribution Our dataset contains a balanced mix of different student… See the full description on the dataset page: https://huggingface.co/datasets/Orib24/Roomly-Student-Bios-Multimodal.texttext-generationn<1K0 likes5 downloads9mo agoHugging Face25Dumoura /amv_genre_multimodal_dataset Estrutura do Conjunto de Dados O conjunto de dados, em formato CSV, foi preparado usando técnicas de visualização de mídia (media visualization). Ele é composto por vídeos de humor sobre temas sociopolíticos, coletados do YouTube. O arquivo contém diversas colunas para uma análise multimodal: amv_genre: O gênero do vídeo, categorizado como drama ou action. Características Visuais: Métricas de cor e brilho, como a média, mediana, desvio padrão e frequência dominante para matiz… See the full description on the dataset page: https://huggingface.co/datasets/Dumoura/amv_genre_multimodal_dataset.tabularfeature-extractionn<1K0 likes4 downloads1y agoHugging Face26DreamTraveler /MultiModal-Planning-Benchmarktext1K<n<10K0 likes3 downloads2y agoHugging Face27Tasfiya025 /Multi-Modal_Sentiment_Analysis_in_E-commercetabularn<1K0 likes3 downloads9mo agoHugging Face28ariefansclub /humanoid-multimodal-intent-resolution Multimodal Intent Resolution Dataset Maps multi-sensor inputs into unified humanoid intents. textn<1K0 likes3 downloads8mo agoHugging Face29Josie1204 /Multimodal-Infomaxtabular1K<n<10K0 likes2 downloads2y agoHugging Face30Shoriful025 /image_caption_pairs_for_multimodalimagen<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.