datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.mind2web_multimodal_test_domain
Dataset Card for "Cross-Domain" Test Split in Multimodal Mind2Web
Note: This dataset is the test split of the Cross-Domain dataset introduced in the paper.
This is a FiftyOne dataset with 4050 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_domain.Multimodal-HC
Multimodal Healthy Control Dataset
100 healthy control volunteers (80 train, 20 test). The test set is released following the Big Cross-Modal Attenuation Correction Challenge (BIC-MAC) held in conjunction with MICCAI 2026.
🌐 PET readouts: hedypet.depict.dk
💻 GitHub: github.com/DEPICT-RH/Multimodal-HC
Citation
If you use this dataset, please cite:
Hinge, C., Høi-Hansen, F.E., Hansen, R.H. et al. A multimodal total-body dynamic [18F]FDG PET/CT/MRI dataset of… See the full description on the dataset page: https://huggingface.co/datasets/DEPICT-RH/Multimodal-HC.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.hard-intersection-multimodal-sample
Hard Intersection Multimodal Samples
Release Notes
Release
Description
v1.0.0
Initial public release.
v1.1.0
Added Unreal Engine assets.Fixed issues in the OpenDRIVE map data.Updated the README to improve documentation and usability.
Dataset Summary
Hard Intersection Multimodal Samples is a curated multimodal dataset of accident-prone urban intersection in Japan for autonomous driving research and development.It provides… See the full description on the dataset page: https://huggingface.co/datasets/dynamic-maps/hard-intersection-multimodal-sample.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans
Multimodal Chest X ray dataset for Normal and Bacterial Pneumonia in Africans | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: imagefolder - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans.pd-voice-full-multimodal-dataset
Parkinson Voice — Full Multimodal Dataset
Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE).
Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel.
Contents
Path
Description
audio/
Waveform clips (healthy / parkinsons), 1134 files
images/mel/
Mel spectrograms
images/spectrogram/
Linear spectrograms
images/mfcc/
MFCC maps
images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.multimodal-LLMs-See-Sentiment
MLLMsent — datasets and experiment results
Every input and every output of "Multimodal LLMs See Sentiment"
(arXiv:2508.16873): the image descriptions generated by six multimodal
LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete
per-fold results of all 141 experiments.
Paper: arXiv:2508.16873
Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.xraydar-multimodal
X-Raydar Multimodal Chest X-Ray Dataset
A multimodal dataset of 979 chest X-ray examinations, each with:
Chest X-ray image (full-resolution PNG, anonymised)
Consensus image-level labels (37 radiological findings, agreed by two expert radiologists with adjudication)
Bounding box annotations on the image from each annotator independently (localising each finding)
Original radiology report text
Report span annotations (token-level labels across 45 categories)
This dataset combines… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/xraydar-multimodal.multimodal-shapes-subset
Dataset Card for multimodal_shapes_subset
This is a grouped multimodal FiftyOne dataset with 2000 samples, each consisting of an rgb image and LiDAR data of objects.
The samples are labeled as cube or sphere. The classes are perfectly balanced (1000 cubes, 1000 spheres).
This dataset is used in this project on github: https://github.com/MatthiasCr/Computer-Vision-Assignment-2/tree/main.
There it is also explained how to use this dataset, visualize it in fiftyone, and convert… See the full description on the dataset page: https://huggingface.co/datasets/MatthiasCr/multimodal-shapes-subset.2026-24679-HW1-Multimodal-Original
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal-Original
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal-Original.Uddessho-Bangla-Multimodal-Intent-Classification
📊 Uddessho Dataset — Multimodal Author Intent Classification
Uddessho (meaning "Intent" in English) is a multimodal dataset created for author intent classification in the low-resource Bangla language.It contains 3,048 social media posts (text + images) labeled into six distinct intent types.
🏷️ Intent Categories & Label Mapping
Label ID
Class Name
0
Advocative
1
Controversial
2
Exhibitionist
3
Expressive
4
Informative
5
Promotive
📂… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Uddessho-Bangla-Multimodal-Intent-Classification.hw1-paper-triage-multimodal
HW1 Multimodal Paper Triage
Purpose
This dataset supports a classroom exercise in assembling and augmenting multimodal data for a personalized research-paper triage system.
Composition and splits
The dataset begins with 100 original multimodal samples. The original samples were split before augmentation using a fixed random seed and stratification by the binary target.
train: 10,070 samples consisting of 70 training originals and 10,000 augmented… See the full description on the dataset page: https://huggingface.co/datasets/ishaanamahajan/hw1-paper-triage-multimodal.2026-24679-HW1-Multimodal
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression experiments.… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal.indonesian-gambling-multimodal-dataset
Indonesian Online Gambling Promotion Multimodal Dataset
This dataset is intended for machine learning research on detecting and filtering online gambling promotions on Indonesian social media. It is a research artifact for content moderation and harmful content detection, not a gambling website or promotional resource.
Research Context
This dataset is created for academic research in harmful content detection and multimodal machine learning. It is intended to help… See the full description on the dataset page: https://huggingface.co/datasets/0xRafie/indonesian-gambling-multimodal-dataset.diabetic-retinopathy-multimodal-progression-africa
DR-Progression — Multimodal 5-Year Diabetic-Retinopathy Risk with Africa-Grounded Synthetic EHR
A multimodal dataset for predicting 5-year diabetic-retinopathy progression and
risk from a fundus image combined with a full systemic electronic health record
(EHR): glycaemic control, diabetes duration, blood pressure, renal function,
comorbidities, treatment, and access-to-care.
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset
family (see also dr-grading and… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-multimodal-progression-africa.CMDS_Multimodal_Document
Dataset Card for Cyrillic Multimodel Document (CMDS)
This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/Cesimbingol54/omega-multimodal.Multimodal_Atrocity_Identification_Dataset
LemkinAI Multimodal Atrocity Identification Dataset
Dataset Overview
This multimodal dataset contains comprehensive documentation of mass atrocities and human rights violations spanning 1980-2025, with 6.8+ million anonymized incident records from 195+ countries. The dataset combines textual documentation with satellite imagery and visual evidence for AI/ML research in atrocity detection, documentation, and prevention systems.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/LemkinAI/Multimodal_Atrocity_Identification_Dataset.diabetic-retinopathy-multimodal-progression-africa
DR-Progression — Multimodal 5-Year Diabetic-Retinopathy Risk with Africa-Grounded Synthetic EHR
A multimodal dataset for predicting 5-year diabetic-retinopathy progression and
risk from a fundus image combined with a full systemic electronic health record
(EHR): glycaemic control, diabetes duration, blood pressure, renal function,
comorbidities, treatment, and access-to-care.
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset
family (see also dr-grading and… See the full description on the dataset page: https://huggingface.co/datasets/arkozaini/diabetic-retinopathy-multimodal-progression-africa.Medical-Multimodal-EN-TH
HealthGPTVL-Translation Medical-Multimodal-EN-TH
This dataset is a bilingual (English-Thai) medical multimodal evaluation dataset containing medical images with corresponding question-answer pairs for visual question answering and translation tasks.
Dataset Details
Dataset Description
This dataset contains 17,047 medical image-text pairs designed for multimodal medical AI evaluation. It includes medical images from various imaging modalities (MRI, CT, X-Ray… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-Multimodal-EN-TH.multimodal_satire
Dataset card for "multimodal_satire"
This is the dataset for the paper A Multi-Modal Method for Satire Detection using Textual and Visual Cues. To obtain the full-text body of the articles, you need to scrape websites using the provided links in the dataset.
GitHub repository: https://github.com/lilyli2004/satire
Reference
If you use this dataset, please cite the following paper:
@inproceedings{li-etal-2020-multi-modal,
title = "A Multi-Modal Method for Satire… See the full description on the dataset page: https://huggingface.co/datasets/phosseini/multimodal_satire.greek-twitter-multimodal-datasetThis is a dataset for sentiment analysis created by text-image pairs collected from greek Twitter.posted: from April 2023 to September 2023Context: general purpose, mostly politics and athleticsTotal pairs: 260The purpose of the dataset is to be used as a test dataset, not for training/fine-tuning a model.Labels: Negative, Neutral, Positivelabelling.xlsx contains the labels for texts,images and both modalities.
Multimodal-Sea-Land
Mulitmodal Sea-Land Segmentation Dataset
Dataset Overview
This dataset contains optical and SAR images from Sentinel-2 and Sentinel-1.
Optical image
The Optical images were acquired from the Sentinel-2 satellite.
SAR image
The SAR images were acquired from the Sentinel-1 satellite. Specifically, the Interferometric Wide Swath (IW) mode was used, incorporating both Vertical-Vertical (VV) and Vertical-Horizontal (VH) polarizations.The Sentinel-1 data… See the full description on the dataset page: https://huggingface.co/datasets/cuibinge/Multimodal-Sea-Land.
