datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multimodal_Fish_Feeding_IntensityMultimodalMathBenchmarks
MultimodalMathBenchmarks
This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026).
It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs.
Canonical Upload Manifest
HF path
Local source
Count
Purpose
SharedMultimodalGrid.csv
SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.pd-voice-full-multimodal-dataset
Parkinson Voice — Full Multimodal Dataset
Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE).
Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel.
Contents
Path
Description
audio/
Waveform clips (healthy / parkinsons), 1134 files
images/mel/
Mel spectrograms
images/spectrogram/
Linear spectrograms
images/mfcc/
MFCC maps
images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.Silver-Multimodal-Dataset
Dataset Overview
The dataset is designed to support the development of machine learning models for detecting daily activities, violence, and fall down scenarios from combined audio and video sources.
The preprocessing pipeline leverages audio feature extraction, human keypoint detection, and relative positional encoding to generate a unified representation for training and inference.
Classes:
0: Daily - Normal indoor activities
1: Violence - Aggressive behaviors
2: Fall Down -… See the full description on the dataset page: https://huggingface.co/datasets/SilverAvocado/Silver-Multimodal-Dataset.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval; use… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.chicken-health-behavior-multimodal
Chicken Health and Behavior Multimodal Dataset (chicken-health-behavior-multimodal)
This dataset provides a comprehensive collection of visual and audio data from chicken farms, specifically designed for early detection of chicken health issues and anomalous behaviors. It aims to support the development of intelligent monitoring systems that can mitigate significant economic losses caused by poultry disease, with a particular focus on future applications.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/IceKhoffi/chicken-health-behavior-multimodal.grandstaff-grandstaff-multimodalegocentric-vr-capture-1h-multimodal-sample
Egocentric VR Capture — 1-Hour Multimodal Inspection Sample
13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.Intel_Robotic_Welding_Multimodal_Dataset
Dataset Card for the Intel Robotic Welding Multimodal Dataset
This dataset was collected to enable multimodal welding defect detection research. The dataset contains over 4000 annotated samples and was collected in an automotive production floor setting in collaboration with a supplier with access to such facilities. Each sample contains a video, associated audio, a time-series from welding sensors, and five post-weld images for a particular weld. A separately licensed… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/Intel_Robotic_Welding_Multimodal_Dataset.Multimodal-Yue-Benchmark
Multimodal Yue Benchmark
Cantonese audio + text benchmark derived from BillBao/Yue-Benchmark (Yue-GSM8K & Yue-MMLU-style tasks). We kept the same task content in Cantonese and added TTS for three Cantonese speakers (hiugaai, hiumaan, wanlung).
Subsets & splits
Config name
Task
Speaker
Splits
mmlu_*
multiple-choice (MMLU-style)
per speaker
train, test
gsm8k_*
math word problems (GSM8K-style)
per speaker
train, test
Example:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/J017athan/Multimodal-Yue-Benchmark.marine-animals-multimodalmultimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.chicken-health-behavior-multimodal
Chicken Health and Behavior Multimodal Dataset (chicken-health-behavior-multimodal)
This dataset provides a comprehensive collection of visual and audio data from chicken farms, specifically designed for early detection of chicken health issues and anomalous behaviors. It aims to support the development of intelligent monitoring systems that can mitigate significant economic losses caused by poultry disease, with a particular focus on future applications.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/Danetuk/chicken-health-behavior-multimodal.keep-it-simple-multimodal
keep-it-simple-multimodal
A mini, standalone multimodal dataset: image+caption, audio+caption, video+caption, lidar, IMU, and
optimal-control state/action pairs. Companion to keep-it-simple
(text), built to feed KairosPretrainingDataset in kairos.
Structure
One generic schema for every row — no per-modality columns, no fixed shape/dtype assumptions:
Column
Type
Description
modality
string
image_caption | audio_caption | video_caption | lidar | imu |… See the full description on the dataset page: https://huggingface.co/datasets/ffurfaro/keep-it-simple-multimodal.multimodal-expert-instruction-samples
Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video
A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside.
▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection
Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.chopin-grandstaff-multimodalmultimodal-genshin-impact
Genshin Impact Fandom Wiki Multimodal Dataset
Github repo here
Description
This dataset is a comprehensive collection of 22,162 fandom wiki pages for the popular game Genshin Impact.
The dataset includes markdown-formatted English content from the wiki, featuring interleaved text, as well as image, video, and audio file links. Additionally, the associated multimodal files (images, videos, and audio) have been downloaded and organized to facilitate the multimodal dataset… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/multimodal-genshin-impact.Geoff_Bremner_Multimodal_Music_Corpus_SAMPLE
Geoff Bremner Multimodal Music Corpus — Sample Release
This is a single-track sample from the Geoff Bremner Multimodal Music Corpus,
a growing, research-grade, commercially licensable dataset of 100% original
music — written, recorded, and produced entirely by one artist .
If this sample meets your needs - please contact me directly for more
Geoff Bremner
https://linktr.ee/gbaudio
License
This dataset is released under CC BY-NC 4.0… See the full description on the dataset page: https://huggingface.co/datasets/geoffbremneraudio/Geoff_Bremner_Multimodal_Music_Corpus_SAMPLE.MultiModalDataset
Dataset Card for MultiModal Dataset
Dataset Description
Dataset Summary
MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation.
The dataset is organized into three subsets:
fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.ClArTTS-multimodalMoroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.multimodal-guide-using-huggingfacemozgach_multimodal_extraMultiModalInstructionFollowingWanJuanSiLu-Multimodal-5Languages
WanJuan·SiLu Multimodal Multilingual Corpus
🌏Dataset Introduction
The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements:
The number of languages has been significantly expanded: Based on the five open-source languages of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-5Languages.kene_multimodal_gift
Kene Multimodal Gift Dataset (Enhanced with Ethnic Languages)
Описание
Мультимодальный духовный датасет с ИКАРОС на испанском, Джив Джаго на хинди и языками народностей России, СНГ и Украины.
Обновления
✅ Добавлены ИКАРОС на испанском языке
✅ Добавлен Джив Джаго на хинди
✅ НОВОЕ: Добавлены языки народностей России, СНГ и Украины
✅ Улучшены мультимодальные данные
✅ Расширена поддержка языков до 50 примеров
Языки
Духовные языки
Русский:… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/kene_multimodal_gift.multimodal-ai-taxonomy
Multimodal AI Taxonomy
A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities.
Dataset Description
This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to:
Navigate the complex landscape of multimodal AI models
Filter models by specific input/output modality combinations
Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.EmotionRecognition_MultimodalEmotionlinesDataset
Dataset Card for "emotion_recognition_multimodal_emotionlines_dataset"
More Information needed
hummel-grandstaff-multimodal
