datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.multiview-pouring
MultiView Pouring Dataset, v1.0.
by Pierre Sermanet, Corey Lynch, Jasmine Hsu and Eric Jang
License
This data is licensed by Google Inc. under a Creative Commons Attribution 4.0 International License.
Downloading
Because of some downloading issues for a specific file, the file was split in two, call https://huggingface.co/datasets/sermanet/multiview-pouring/blob/main/tfrecords/test/whiteorange_to_clear1_real_combining.sh to recombine the parts.… See the full description on the dataset page: https://huggingface.co/datasets/sermanet/multiview-pouring.toon-multimedia-datastanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.aloha_multiviewhard-intersection-multimodal-sample
Dataset Card for Hard Intersection Multimodal Sample
Dataset Details
Dataset Description
Hard Intersection Multimodal Sample is a curated multimodal dataset of an accident-prone six-way urban intersection in Tokyo, Japan (Takanawadai) captured with an industrial mobile mapping system. The dataset provides synchronized multi-camera views, LiDAR point clouds, vehicle trajectories, HD maps in multiple formats, and semantic annotations for autonomous… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/hard-intersection-multimodal-sample.dvla-multi10x5-orig-25fpsEgoDex-PickPlace-YAM-14dof-multiview
EgoDex → YAM 14-DOF, Multiview (LeRobot v2.1)
Egocentric human hand-manipulation demonstrations from EgoDex retargeted to a
YAM bimanual robot (14-DOF), packaged as a LeRobot v2.1 dataset with three
synthesized camera views. The observation schema is drop-in compatible with
angkul07/abc-teleop for
cotraining (identical Hz, camera keys, and action convention).
At a glance
Episodes
8,842
Frames
1,074,893
Control rate
30 Hz (30 fps video)
Robot
YAM… See the full description on the dataset page: https://huggingface.co/datasets/angkul07/EgoDex-PickPlace-YAM-14dof-multiview.G1_WBT_Brainco_Pick_Up_Multiple_CushionsMultiCamWarpNEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.smolvla_multiblockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 10,
"total_frames": 4437,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/smolvla_multiblock.frame-synced-multiplayer
Origin Lab Frame-Synced Multiplayer: Eight Players, One Frame Clock
Six to eight players, each on their own PC on residential internet, each
recording their own live view of one match, and frame k on every machine is
the same server instant, verified four independent ways, with the
verification script in this repo. Every player ships the full engine stack
at 1080p / 60 FPS on one shared frame grid. Because the alignment is measured
rather than assumed, the release also… See the full description on the dataset page: https://huggingface.co/datasets/originlab/frame-synced-multiplayer.multivsr
Dataset: MultiVSR
We introduce a large-scale multilingual lip-reading dataset: MultiVSR. The dataset comprises a total of 12,000 hours of video footage, covering English + 12 non-English languages. MultiVSR is a massive dataset with a huge diversity in terms of the speakers as well as languages, with approximately 1.6M video clips across 123K YouTube videos. Please check the website for samples.
Download instructions
Please check the GitHub repo to download… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/multivsr.test_multiview_3d_reconstructionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 757,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/test_multiview_3d_reconstruction.test_multiview_3d_reconstruction_3camsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 375,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/test_multiview_3d_reconstruction_3cams.G1_WBT_Inspire_Pick_Up_Multiple_Cushionsanimeshooter-multishot-iclora
AnimeShooter → LTX-2.3 IC-LoRA multi-shot training pairs
276 reference-guided multi-shot training pairs derived from the
AnimeShooter dataset
(Qiu et al., 2025, arXiv:2506.03126), formatted for the
ltx-community/ltx2-lora-trainer
Space in IC-LoRA mode.
Each sample is:
videos/<id>.mp4 — a ~5.2s clip (126 frames, H.264) spanning 2–3 consecutive shots
(real cuts inside the clip) from an animated YouTube video listed in AnimeShooter.
videos/<id>_reference.png — a 768×512… See the full description on the dataset page: https://huggingface.co/datasets/linoyts/animeshooter-multishot-iclora.Silver-Multimodal-Dataset
Dataset Overview
The dataset is designed to support the development of machine learning models for detecting daily activities, violence, and fall down scenarios from combined audio and video sources.
The preprocessing pipeline leverages audio feature extraction, human keypoint detection, and relative positional encoding to generate a unified representation for training and inference.
Classes:
0: Daily - Normal indoor activities
1: Violence - Aggressive behaviors
2: Fall Down -… See the full description on the dataset page: https://huggingface.co/datasets/SilverAvocado/Silver-Multimodal-Dataset.Multisence360_Dataset
MultiScene360 Dataset
A Real-World Multi-Camera Video Dataset for Generative Vision AI
📌 Overview
The MultiScene360 Dataset is designed to advance generative vision AI by providing synchronized multi-camera footage from real-world environments.
💡 Key Applications:✔ Video generation & view synthesis✔ 3D reconstruction & neural rendering✔ Digital human animation systems✔ Virtual/augmented reality development
📊 Dataset Specifications (Public Version)… See the full description on the dataset page: https://huggingface.co/datasets/Eric-maadaa/Multisence360_Dataset.so101_multi_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 9,
"total_frames": 3131,
"total_tasks": 1,
"total_videos": 18,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/fbeltrao/so101_multi_task.omy_f3m_multi_spacemouse_dual_arm_pin_insertionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omy_f3m_multi_spacemouse",
"total_episodes": 52,
"total_frames": 35275,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/omy_f3m_multi_spacemouse_dual_arm_pin_insertion.E-VLA-MultiLightmulti-perspective-dataset-plain-zebras
Dataset Card for Multi-Perspective Dataset of Plains Zebras
Four autonomous swarm missions over plains zebras, recorded simultaneously from three or four drones, with every camera and flight-log clock placed on one timebase.
Dataset Details
This dataset covers four autonomous drone-swarm missions monitoring plains zebras
(Equus quagga) at Ol Pejeta Conservancy, Laikipia County, Kenya, flown between
21 and 24 February 2026. Each mission flew one scout at high… See the full description on the dataset page: https://huggingface.co/datasets/edouard-rolland/multi-perspective-dataset-plain-zebras.insert-usb-ethernet-multipin-0901This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 100,
"total_frames": 110773,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/insert-usb-ethernet-multipin-0901.stack_multi_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 51,
"total_frames": 44895,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/danielsanjosepro/stack_multi_v1.cs2-v3-prompt-comparison-7-examples-with-multiaction-gemini35
CS2 V3 七案例 Gemini 3.5 Flash 最终结果对比 / Seven-case Gemini 3.5 Flash Comparison
本 README 展示 Gemini 3.5 Flash 对同一批 7 个视频的最终打标结果:每个案例先显示视频,再用左右两列并排展示两轮和七轮的完整 English JSON 与中文 JSON;内容直接展开,字号保持较小以便对照。
This README shows Gemini 3.5 Flash final labels for the same 7 videos. Each case places the video first, then displays complete English and Chinese JSON side by side: two-round on the left and seven-round on the right.
两轮与七轮的 API 输入详情通过顶部索引查看;中文侧保持与英文 JSON 相同的键、时间边界、数组长度和 Action 标签。
API… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-prompt-comparison-7-examples-with-multiaction-gemini35.omy_f3m_multi_Disconnect-0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_multi",
"total_episodes": 100,
"total_frames": 61744,
"total_tasks": 1,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/omy_f3m_multi_Disconnect-0.multiplied_so101-table-cleanupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 1358,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CrawlAiHuggingFace/multiplied_so101-table-cleanup.
