datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.plasticc---
description: 'The Photometric LSST Astronomical Time-Series Classification Challenge
(PLAsTiCC) is a community-wide challenge to spur development of algorithms to classify
astronomical transients. The Large Synoptic Survey Telescope (LSST) will discover
tens of thousands of transient phenomena every single night. To deal with this massive
onset of data, automated algorithms to classify and sort astronomical transients
are crucial.
'
homepage: https://zenodo.org/records/2539456… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/plasticc.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval; use… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.stanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.multi-modal-derived-brain-network
PPMI Connectivity Graphs — HF Staging (Derivatives)
This dataset ships ready-to-use functional brain connectivity graphs derived from the PPMI cohort in a BIDS-ish derivatives layout. For each subject and parcellation, we include:
ROI time-series (*_desc-timeseries_parc-<name>.mat)
Pearson correlation connectivity matrix (*_desc-correlation_matrix_parc-<name>.mat)
JSON sidecars with summary fields (nodes, measure, symmetric/weighted flags)
Contents
data/… See the full description on the dataset page: https://huggingface.co/datasets/pakkinlau/multi-modal-derived-brain-network.gaia---
description: 'Spectral (BP/RP), photometric, and astrometric dataset based on Gaia
DR3.
'
homepage: https://www.cosmos.esa.int/web/gaia/dr3
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% If you have used Gaia DR3 data in your research, \ please use the following acknowledgement:\n% \n% This work has made use of data \ from the European Space Agency (ESA) mission\n% {\it Gaia} (\url{https://www.cosmos.esa.int/gaia}),\
\ processed by the {\it Gaia}\n% Data Processing and Analysis… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/gaia.jwst---
description: 'Image dataset based on a combination of JWST deep fields from DJA: CEERS,
NGDEEP, JADES, PRIMER
'
homepage: https://dawn-cph.github.io/dja/index.html
version: 1.1.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://dawn-cph.github.io/dja/index.html\n\
% We kindly request all scientific papers based on data or products downloaded from \ the Dawn JWST Archive (DJA) to include the following acknowledgement:\n% \n% (Some \ of) The data products presented herein were… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/jwst.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.egocentric-vr-capture-1h-multimodal-sample
Egocentric VR Capture — 1-Hour Multimodal Inspection Sample
13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.VQAv2_test_no_image
Dataset Card for "VQAv2_test_no_image"
More Information needed
multimodal_qa_dataset_v4_trainhsc---
description: 'Image dataset based on HSC SSP PRD3.
'
homepage: https://hsc-release.mtk.nao.ac.jp/doc/
version: 1.0.0
citation: "% CITATION\n@article{Aihara_2017,\n title={The Hyper Suprime-Cam SSP \ Survey: Overview and survey design},\n volume={70},\n ISSN={2053-051X},\n \ url={http://dx.doi.org/10.1093/pasj/psx066},\n DOI={10.1093/pasj/psx066},\n \ number={SP1},\n journal={Publications of the Astronomical Society of Japan},\n \ publisher={Oxford University Press… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/hsc.multimodal_qa_dataset_v3_trainmultimodal-lucas
Dataset card for Multi-modal LUCAS
Dataset summary
Multi-modal LUCAS aims at being a curated vision-language dataset from LUCAS survey data and in-situ field photos. LUCAS (Land Use/Cover Area Frame statistical Survey) is a land-monitoring exercise conducted by EUROSTAT in close cooperation with the Directorate-General responsible for Agriculture, with technical support from the Joint Research Centre (JRC). The survey has been repeated every three years since 2006… See the full description on the dataset page: https://huggingface.co/datasets/stemauro/multimodal-lucas.stanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
meta/info.json
{
"codebase_version": "v2.0",
"data_path": "data/chunk-{episode_chunk:03d}/train-{episode_index:05d}-of-{total_episodes:05d}.parquet",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"keys": [
"observation.state"… See the full description on the dataset page: https://huggingface.co/datasets/aliberts/stanford_kuka_multimodal_dataset.ego-multimodal
ego-multimodal: Full Body Motion Capture with Finger Dexterity and General Motion Retargeting (GMR)
Research Use Only — This dataset is released under CC-BY-NC-4.0 and is intended strictly for non-commercial research purposes. Commercial use is prohibited.
A full body motion capture dataset with finger dexterity, recorded with MoWare (10 IMU sensors — 5 upper body, 5 lower body) and the Phi9 Glove for fine-grained finger tracking. This demo uses upper body sensors and the Phi9… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/ego-multimodal.sdss---
description: 'Spectra dataset based on SDSS-IV.
'
homepage: https://www.sdss.org/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\
% \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for High-Performance Computing at the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/sdss.multimodality-poc-llama31-ruler16k
Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K
Raw pre-RoPE query and hidden-state tensors captured during prefill, used
to study whether the per-(layer, kv_head) query distribution is unimodal
Gaussian (the assumption underpinning Expected Attention's MGF closed-form
in kvpress).
What's in here
65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5
prompts).
Each file (~414 MB) contains:
field
dtype
shape
meaning
hidden
float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.multi-modal-peg-in-square-hole-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 51,
"total_frames": 8074,
"total_tasks": 1,
"total_videos": 153,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-test.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.agent-spaces-tracesmultimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.desi---
description: 'Spectra dataset based on DESI EDR SV3.
'
homepage: https://data.desi.lbl.gov/doc
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% From https://data.desi.lbl.gov/doc/acknowledgments/\
\ : \n% \n% The Dark Energy Spectroscopic Instrument (DESI) data are licensed under \ the Creative Commons Attribution 4.0 International License (“CC BY 4.0”, Summary, \ Full Legal Code). Users are free to share, copy, redistribute, adapt, transform \ and build upon the DESI data… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/desi.
