datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
plasticc---
description: 'The Photometric LSST Astronomical Time-Series Classification Challenge
(PLAsTiCC) is a community-wide challenge to spur development of algorithms to classify
astronomical transients. The Large Synoptic Survey Telescope (LSST) will discover
tens of thousands of transient phenomena every single night. To deal with this massive
onset of data, automated algorithms to classify and sort astronomical transients
are crucial.
'
homepage: https://zenodo.org/records/2539456… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/plasticc.stanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.multi-modal-derived-brain-network
PPMI Connectivity Graphs — HF Staging (Derivatives)
This dataset ships ready-to-use functional brain connectivity graphs derived from the PPMI cohort in a BIDS-ish derivatives layout. For each subject and parcellation, we include:
ROI time-series (*_desc-timeseries_parc-<name>.mat)
Pearson correlation connectivity matrix (*_desc-correlation_matrix_parc-<name>.mat)
JSON sidecars with summary fields (nodes, measure, symmetric/weighted flags)
Contents
data/… See the full description on the dataset page: https://huggingface.co/datasets/pakkinlau/multi-modal-derived-brain-network.gaia---
description: 'Spectral (BP/RP), photometric, and astrometric dataset based on Gaia
DR3.
'
homepage: https://www.cosmos.esa.int/web/gaia/dr3
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% If you have used Gaia DR3 data in your research, \ please use the following acknowledgement:\n% \n% This work has made use of data \ from the European Space Agency (ESA) mission\n% {\it Gaia} (\url{https://www.cosmos.esa.int/gaia}),\
\ processed by the {\it Gaia}\n% Data Processing and Analysis… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/gaia.egocentric-vr-capture-1h-multimodal-sample
Egocentric VR Capture — 1-Hour Multimodal Inspection Sample
13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.hsc---
description: 'Image dataset based on HSC SSP PRD3.
'
homepage: https://hsc-release.mtk.nao.ac.jp/doc/
version: 1.0.0
citation: "% CITATION\n@article{Aihara_2017,\n title={The Hyper Suprime-Cam SSP \ Survey: Overview and survey design},\n volume={70},\n ISSN={2053-051X},\n \ url={http://dx.doi.org/10.1093/pasj/psx066},\n DOI={10.1093/pasj/psx066},\n \ number={SP1},\n journal={Publications of the Astronomical Society of Japan},\n \ publisher={Oxford University Press… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/hsc.VQAv2_test_no_image
Dataset Card for "VQAv2_test_no_image"
More Information needed
jwst---
description: 'Image dataset based on a combination of JWST deep fields from DJA: CEERS,
NGDEEP, JADES, PRIMER
'
homepage: https://dawn-cph.github.io/dja/index.html
version: 1.1.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://dawn-cph.github.io/dja/index.html\n\
% We kindly request all scientific papers based on data or products downloaded from \ the Dawn JWST Archive (DJA) to include the following acknowledgement:\n% \n% (Some \ of) The data products presented herein were… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/jwst.ego-multimodal
ego-multimodal: Full Body Motion Capture with Finger Dexterity and General Motion Retargeting (GMR)
Research Use Only — This dataset is released under CC-BY-NC-4.0 and is intended strictly for non-commercial research purposes. Commercial use is prohibited.
A full body motion capture dataset with finger dexterity, recorded with MoWare (10 IMU sensors — 5 upper body, 5 lower body) and the Phi9 Glove for fine-grained finger tracking. This demo uses upper body sensors and the Phi9… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/ego-multimodal.multimodal_qa_dataset_v4_trainmultimodal_qa_dataset_v3_trainstanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
meta/info.json
{
"codebase_version": "v2.0",
"data_path": "data/chunk-{episode_chunk:03d}/train-{episode_index:05d}-of-{total_episodes:05d}.parquet",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"keys": [
"observation.state"… See the full description on the dataset page: https://huggingface.co/datasets/aliberts/stanford_kuka_multimodal_dataset.sdss---
description: 'Spectra dataset based on SDSS-IV.
'
homepage: https://www.sdss.org/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\
% \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for High-Performance Computing at the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/sdss.multi-modal-peg-in-square-hole-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 51,
"total_frames": 8074,
"total_tasks": 1,
"total_videos": 153,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-test.desi---
description: 'Spectra dataset based on DESI EDR SV3.
'
homepage: https://data.desi.lbl.gov/doc
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% From https://data.desi.lbl.gov/doc/acknowledgments/\
\ : \n% \n% The Dark Energy Spectroscopic Instrument (DESI) data are licensed under \ the Creative Commons Attribution 4.0 International License (“CC BY 4.0”, Summary, \ Full Legal Code). Users are free to share, copy, redistribute, adapt, transform \ and build upon the DESI data… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/desi.VQAv2_validation_no_image
Dataset Card for "VQAv2_validation_no_image"
More Information needed
multi-modal-peg-in-square-hole-image-aloneThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 11,
"total_frames": 1892,
"total_tasks": 1,
"total_videos": 33,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:11"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-image-alone.multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.multimodal-benchmarks-minitess---
description: 'TESS Light Curves From Full Frame Images ("TESS-SPOC")
'
homepage: https://archive.stsci.edu/hlsp/tess-spoc
version: 1.0.0
citation: "% % ACKNOWLEDGEMENT\n% % From: https://archive.stsci.edu/publishing/mission-acknowledgements\n\
% This paper includes data collected with the TESS mission, obtained from the MAST \ data archive at the Space Telescope Science Institute (STScI). Funding for the \ TESS mission is provided by the NASA Explorer Program. STScI is operated by… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/tess.VQAv2_testdev_no_image
Dataset Card for "VQAv2_testdev_no_image"
More Information needed
diabetic-retinopathy-multimodal-progression-africa
DR-Progression — Multimodal 5-Year Diabetic-Retinopathy Risk with Africa-Grounded Synthetic EHR
A multimodal dataset for predicting 5-year diabetic-retinopathy progression and
risk from a fundus image combined with a full systemic electronic health record
(EHR): glycaemic control, diabetes duration, blood pressure, renal function,
comorbidities, treatment, and access-to-care.
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset
family (see also dr-grading and… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-multimodal-progression-africa.multimodal_care
CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions
CARE is a multimodal English dataset comprising approximately 144 hours of clinical interviews from 622 profiles across 12 medical conditions plus a control cohort.
The dataset contains participant metadata together with pre-computed acoustic, linguistic, and visual features extracted from interview recordings. CARE is designed to support research in speech and language… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/multimodal_care.Imagenet1k_trainVQAv2_test_no_image_split_5
Dataset Card for "VQAv2_test_no_image_split_5"
More Information needed
btsbot---
description: 'This is the production version of the BTSbot training set, limited to
public (programid=1) ZTF alerts.
Original codebase: https://github.com/nabeelre/BTSbot
'
homepage: https://zenodo.org/records/10839691
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % Based on the Acknowledgements in Rehemtulla et \ al. (2024). We suggest including a variant of the following in your acknowledgements:\n % A great number of people have contributed to BTS and BTS scanning over the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/btsbot.VQAv2_minival
Dataset Card for "VQAv2_minival"
More Information needed
