datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.multimodal-LLMs-See-Sentiment
MLLMsent — datasets and experiment results
Every input and every output of "Multimodal LLMs See Sentiment"
(arXiv:2508.16873): the image descriptions generated by six multimodal
LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete
per-fold results of all 141 experiments.
Paper: arXiv:2508.16873
Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.2026-24679-HW1-Multimodal-Original
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal-Original
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal-Original.hw1-paper-triage-multimodal
HW1 Multimodal Paper Triage
Purpose
This dataset supports a classroom exercise in assembling and augmenting multimodal data for a personalized research-paper triage system.
Composition and splits
The dataset begins with 100 original multimodal samples. The original samples were split before augmentation using a fixed random seed and stratification by the binary target.
train: 10,070 samples consisting of 70 training originals and 10,000 augmented… See the full description on the dataset page: https://huggingface.co/datasets/ishaanamahajan/hw1-paper-triage-multimodal.2026-24679-HW1-Multimodal
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression experiments.… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal.CMDS_Multimodal_Document
Dataset Card for Cyrillic Multimodel Document (CMDS)
This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.diabetic-retinopathy-multimodal-progression-africa
DR-Progression — Multimodal 5-Year Diabetic-Retinopathy Risk with Africa-Grounded Synthetic EHR
A multimodal dataset for predicting 5-year diabetic-retinopathy progression and
risk from a fundus image combined with a full systemic electronic health record
(EHR): glycaemic control, diabetes duration, blood pressure, renal function,
comorbidities, treatment, and access-to-care.
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset
family (see also dr-grading and… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-multimodal-progression-africa.Medical-Multimodal-EN-TH
HealthGPTVL-Translation Medical-Multimodal-EN-TH
This dataset is a bilingual (English-Thai) medical multimodal evaluation dataset containing medical images with corresponding question-answer pairs for visual question answering and translation tasks.
Dataset Details
Dataset Description
This dataset contains 17,047 medical image-text pairs designed for multimodal medical AI evaluation. It includes medical images from various imaging modalities (MRI, CT, X-Ray… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-Multimodal-EN-TH.diabetic-retinopathy-multimodal-progression-africa
DR-Progression — Multimodal 5-Year Diabetic-Retinopathy Risk with Africa-Grounded Synthetic EHR
A multimodal dataset for predicting 5-year diabetic-retinopathy progression and
risk from a fundus image combined with a full systemic electronic health record
(EHR): glycaemic control, diabetes duration, blood pressure, renal function,
comorbidities, treatment, and access-to-care.
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset
family (see also dr-grading and… See the full description on the dataset page: https://huggingface.co/datasets/arkozaini/diabetic-retinopathy-multimodal-progression-africa.multimodal_satire
Dataset card for "multimodal_satire"
This is the dataset for the paper A Multi-Modal Method for Satire Detection using Textual and Visual Cues. To obtain the full-text body of the articles, you need to scrape websites using the provided links in the dataset.
GitHub repository: https://github.com/lilyli2004/satire
Reference
If you use this dataset, please cite the following paper:
@inproceedings{li-etal-2020-multi-modal,
title = "A Multi-Modal Method for Satire… See the full description on the dataset page: https://huggingface.co/datasets/phosseini/multimodal_satire.
