datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.hard-intersection-multimodal-sample
Dataset Card for Hard Intersection Multimodal Sample
Dataset Details
Dataset Description
Hard Intersection Multimodal Sample is a curated multimodal dataset of an accident-prone six-way urban intersection in Tokyo, Japan (Takanawadai) captured with an industrial mobile mapping system. The dataset provides synchronized multi-camera views, LiDAR point clouds, vehicle trajectories, HD maps in multiple formats, and semantic annotations for autonomous… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/hard-intersection-multimodal-sample.stanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.Silver-Multimodal-Dataset
Dataset Overview
The dataset is designed to support the development of machine learning models for detecting daily activities, violence, and fall down scenarios from combined audio and video sources.
The preprocessing pipeline leverages audio feature extraction, human keypoint detection, and relative positional encoding to generate a unified representation for training and inference.
Classes:
0: Daily - Normal indoor activities
1: Violence - Aggressive behaviors
2: Fall Down -… See the full description on the dataset page: https://huggingface.co/datasets/SilverAvocado/Silver-Multimodal-Dataset.office_multimodal_sweep
Office Multimodal Manipulation Sweep
A dual-arm (ALOHA) manipulation dataset collected in the RoboPRO / RoboTwin
simulator. Every task is run over many scene seeds and four distinct planner
motion modes, so the same (task, seed) is recorded four ways that differ only
in how the arm moves — a controlled source of trajectory-level multimodality.
Contents
20 office tasks × 20 scene seeds × 4 motion modes = 1,600 episodes.
Outcome tally: 919 mode_success · 646… See the full description on the dataset page: https://huggingface.co/datasets/choijoshua16/office_multimodal_sweep.egocentric-vr-capture-1h-multimodal-sample
Egocentric VR Capture — 1-Hour Multimodal Inspection Sample
13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.Intel_Robotic_Welding_Multimodal_Dataset
Dataset Card for the Intel Robotic Welding Multimodal Dataset
This dataset was collected to enable multimodal welding defect detection research. The dataset contains over 4000 annotated samples and was collected in an automotive production floor setting in collaboration with a supplier with access to such facilities. Each sample contains a video, associated audio, a time-series from welding sensors, and five post-weld images for a particular weld. A separately licensed… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/Intel_Robotic_Welding_Multimodal_Dataset.egocentric-video-imu-multimodal-sample-v1
Origin Data Lab — Egocentric Video + 6-DoF IMU Sample
This public technical sample demonstrates Origin Data Lab's capability to collect, structure, quality-check, and package real-world egocentric multimodal data for AI, robotics, and embodied-AI applications.
The sample contains egocentric video, task-level 6-DoF IMU data, sanitized metadata, and machine-measured sensor QC.
This repository is a limited public capability sample, not a complete production dataset.… See the full description on the dataset page: https://huggingface.co/datasets/origindatalab/egocentric-video-imu-multimodal-sample-v1.ego-multimodal
ego-multimodal: Full Body Motion Capture with Finger Dexterity and General Motion Retargeting (GMR)
Research Use Only — This dataset is released under CC-BY-NC-4.0 and is intended strictly for non-commercial research purposes. Commercial use is prohibited.
A full body motion capture dataset with finger dexterity, recorded with MoWare (10 IMU sensors — 5 upper body, 5 lower body) and the Phi9 Glove for fine-grained finger tracking. This demo uses upper body sensors and the Phi9… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/ego-multimodal.multi-modal-peg-in-square-hole-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 51,
"total_frames": 8074,
"total_tasks": 1,
"total_videos": 153,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-test.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.multimodal-expert-instruction-samples
Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video
A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside.
▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection
Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.keep-it-simple-multimodal
keep-it-simple-multimodal
A mini, standalone multimodal dataset: image+caption, audio+caption, video+caption, lidar, IMU, and
optimal-control state/action pairs. Companion to keep-it-simple
(text), built to feed KairosPretrainingDataset in kairos.
Structure
One generic schema for every row — no per-modality columns, no fixed shape/dtype assumptions:
Column
Type
Description
modality
string
image_caption | audio_caption | video_caption | lidar | imu |… See the full description on the dataset page: https://huggingface.co/datasets/ffurfaro/keep-it-simple-multimodal.multi-modal-peg-in-square-hole-image-aloneThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 11,
"total_frames": 1892,
"total_tasks": 1,
"total_videos": 33,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:11"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-image-alone.multimodal_care
CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions
CARE is a multimodal English dataset comprising approximately 144 hours of clinical interviews from 622 profiles across 12 medical conditions plus a control cohort.
The dataset contains participant metadata together with pre-computed acoustic, linguistic, and visual features extracted from interview recordings. CARE is designed to support research in speech and language… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/multimodal_care.YT-DemTalk
YouTube Dementia Speaking (URLs + Labels)
Last updated: 2025-10-30
A large dataset of YouTube links where a single person speaks to camera, with labels that correspond to the self reported dementia/alzheimer disgnosis in the table (e.g., dementia vs. neurotypical).The repository intentionally stores only links and annotations, not the videos themselves.
Files
data/train.csv — canonical CSV (recommended)
Columns
label: (describe this column)
url: (describe… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-multimodal-researcher/YT-DemTalk.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/Cesimbingol54/omega-multimodal.multimodal-oil-gas-benchmark
Dataset Card for Dataset Name
This is a dataset to reproduce our work, "A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection", accepted to NeurIPS 2025 Benchmarks and Datasets Track.
Dataset Details
Dataset Sources [optional]
Repository: https://github.com/climate-nlp/multimodal-oil-gas-benchmark
Paper: A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection (NeurIPS 2025)… See the full description on the dataset page: https://huggingface.co/datasets/climate-nlp/multimodal-oil-gas-benchmark.real-world-multimodal-field-sample
Origin Data Lab --- Real-World Multimodal Field Sample
A compact technical sample demonstrating Origin Data Lab's real-world
field data collection and structured data-operations workflow.
Capture → Metadata / IMU → Validation → QC → Privacy Processing →
Structured Delivery
This public repository is intentionally limited in scope. It is designed
to provide verifiable evidence of our workflow without publishing
sensitive raw footage, contributor-identifying information, precise… See the full description on the dataset page: https://huggingface.co/datasets/origindatalab/real-world-multimodal-field-sample.mff-multimodal-dataset
MFF Multimodal Video Editing Dataset
This dataset contains source videos, text editing prompts, and style reference images
used for multimodal video editing experiments.
Each row in metadata.jsonl pairs one source video, one text prompt, and one style
image. The dataset contains 117 rows: 13 videos × 3 prompts × 3 style images.
Columns
example_id: Unique row identifier.
frame_group: Source video group, one of 8-frames, 36-frames, or 90-frames.
num_frames: Number of… See the full description on the dataset page: https://huggingface.co/datasets/AviadDahan/mff-multimodal-dataset.reachy2-kitchen-multimodal
Reachy2 Kitchen Multimodal Dataset
Teleoperation dataset for Reachy2 mobile bi-manipulator performing kitchen tasks in the SAI platform. Collected using sai-zoo teleoperation tool and converted to LeRobot format.
Dataset Statistics
Metric
Value
Episodes
219
Total Frames
23,644
Tasks
9
FPS
20
Camera Resolution
256×256
Tasks
Task
Description
Drawer
Open/close left drawer, Open/close right drawer
Stove
Turn on/off front-left… See the full description on the dataset page: https://huggingface.co/datasets/CompeteSAI/reachy2-kitchen-multimodal.multi_modality_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 6,
"total_frames": 2543,
"total_tasks": 1,
"total_videos": 18,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seeingrain/multi_modality_test.multi-modal-peg-in-square-holeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 201,
"total_frames": 28616,
"total_tasks": 1,
"total_videos": 603,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole.Neurodiverse_Multimodal_Datasetltx-2.3-foley-aotiUMED-Urdu-Multimodal-Emotion-DatasetUrdu-Multimodal-Emotion-DatasetPanchayat-multimodal-datasetThis dataset was created from scratch for research on humour in conversational dialogues.The aim was to contribute a high-quality multimodal resource for the Hindi language.
The dataset contains:
Hindi text written in Devanagari script
English words transcribed exactly as spoken
Conversational multimodal data including text, audio, and video
The transcription style preserves natural conversational context and code-mixing patterns commonly found in spoken Hindi.
Citation
If you… See the full description on the dataset page: https://huggingface.co/datasets/Abhis4e/Panchayat-multimodal-dataset.
