datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.DexJoCo-Datasets-Raw
Dataset Card for DexJoCo
This dataset provides the raw data for DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo.
See more at:
Hugging Face Paper: https://huggingface.co/papers/2605.16257
arXiv Paper: https://arxiv.org/abs/2605.16257
GitHub: https://github.com/brave-eai/dexjoco
BibTeX:
@misc{wang2026dexjocobenchmarktoolkittaskoriented,
title={DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo}… See the full description on the dataset page: https://huggingface.co/datasets/DexJoCo/DexJoCo-Datasets-Raw.benchmark-datasets
Latency-Sensitive Bench datasets
Accepted zero-latency teacher rollouts for the supported benchmark tasks.
Viewer subsets
humanoidbench_balance_simple: 90 training and 10 validation episodes. The
observation.image values are PNG bytes declared as the Hugging Face Image
feature, so the Dataset Viewer renders them instead of showing their encoded
representation. Canonical LeRobot MP4 files remain under each split's
videos/ directory.
mikasa_intercept_grab_fast:… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/benchmark-datasets.DexJoCo-Datasets-LeRobot
Dataset Card for DexJoCo
This dataset provides the LeRobot format dataset for DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo.
See more at:
Hugging Face Paper: https://huggingface.co/papers/2605.16257
arXiv Paper: https://arxiv.org/abs/2605.16257
GitHub: https://github.com/brave-eai/dexjoco
BibTeX:
@misc{wang2026dexjocobenchmarktoolkittaskoriented,
title={DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo}… See the full description on the dataset page: https://huggingface.co/datasets/DexJoCo/DexJoCo-Datasets-LeRobot.TALKtoME
TALKtoME: Educational Materials for Speech and Language Acquisition in Autism
This dataset contains image and video samples of action verbs and verb+noun pairs. It is designed to support machine learning tasks related to visual understanding, action recognition, and language grounding.
Dataset Structure
The dataset contains the following main folders:
action_verbs/images/: image samples organized by action verb categories.
action_verbs/videos/: video samples organized by… See the full description on the dataset page: https://huggingface.co/datasets/LSL-datasets/TALKtoME.solaris-eval-datasets
Solaris Eval Datasets
Project Page | Paper | Github
Evaluation datasets collected via SolarisEngine to evaluate the Solaris multiplayer world model for Minecraft. Refer to Solaris repository for downloading and evaluation running instructions.
You can also use it to benchmark any multi-agent action-conditioned video model.
Dataset Info
The dataset contains videos and actions for two players in Minecraft at 720p and 20 fps.
You can find the action space in the training… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/solaris-eval-datasets.yamlab_datasetsAIC26-Datasetscoffee-vla-datasets-tovgp_datasets
vgp_datasets
Source datasets for video_gen_physics (Video World benchmark).
LeRoBot layout (parquet + mp4):
{embodiment}/{view_type}/{makovian,non_makovian}/
data/chunk-XXX/episode_XXXXXX.parquet
videos/chunk-XXX/{view_name}/episode_XXXXXX.mp4
meta/{info.json, episodes.jsonl, ...}
makovian: next state depends mainly on current state + action
non_makovian: longer temporal dependencies
Canonical trees: bimanual/, humanoid/, single_arm/ (singleview + multiview).
Folders… See the full description on the dataset page: https://huggingface.co/datasets/doanh25032004/vgp_datasets.argus-datasetsvideo_caption_datasetsDexJoCo-Datasets-LeRobot
Dataset Card for DexJoCo
This dataset provides the LeRobot format dataset for DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo.
See more at:
Hugging Face Paper: https://huggingface.co/papers/2605.16257
arXiv Paper: https://arxiv.org/abs/2605.16257
GitHub: https://github.com/brave-eai/dexjoco
BibTeX:
@misc{wang2026dexjocobenchmarktoolkittaskoriented,
title={DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on… See the full description on the dataset page: https://huggingface.co/datasets/zhou-sicheng0731/DexJoCo-Datasets-LeRobot.datasets-cache0529_DATASETSThis dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"total_episodes": 121,
"total_frames": 90808,
"total_videos": 363,
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_tasks": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HSJUSER/0529_DATASETS.Visual-WetlandBirds-Dataset
Dataset Card for Visual WetlandBirds Dataset
The Visual WetlandBirds Dataset is a fine-grained spatio-temporal dataset specifically designed for bird behavior detection and species classification. This version has been converted to work well with the Hugging Face Hub, with the original dataset available at Zenodo.
The dataset was introduced in the paper Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/academic-datasets/Visual-WetlandBirds-Dataset.Franka-Datasets-v2-5kunderwater-img-imu-sonar-datasetsDexJoCo-Datasets-Foresight
Dataset Card for DexJoCo
This dataset provides the LeRobot format dataset for DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo.
See more at:
Hugging Face Paper: https://huggingface.co/papers/2605.16257
arXiv Paper: https://arxiv.org/abs/2605.16257
GitHub: https://github.com/brave-eai/dexjoco
BibTeX:
@misc{wang2026dexjocobenchmarktoolkittaskoriented,
title={DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on… See the full description on the dataset page: https://huggingface.co/datasets/AmurBear/DexJoCo-Datasets-Foresight.arx5_real_world_datasets
arx5_real_world_datasets
This repository contains the real-world robot datasets collected on the ARX-5 platform as presented in the paper Self-Correcting VLA: Online Action Refinement via Sparse World Imagination.
The dataset is part of the SC-VLA framework, which enhances robotic physical grounding through intrinsic self-improvement using Sparse World Imagination (SPI) and Online Action Refinement (OAR).
Code: https://github.com/Kisaragi0/SC-VLA
Paper: Self-Correcting VLA: Online… See the full description on the dataset page: https://huggingface.co/datasets/Kisaragi0/arx5_real_world_datasets.datasetSEGMENTfordraw
DeliverStraw Drawer-Opening Segments
This dataset contains the drawer-opening portion extracted from all 504 target
RoboCasa DeliverStraw demonstrations. It is derived from
target/composite/DeliverStraw/20250813 in NVIDIA's
PhysicalAI-Robotics-Manipulation-Kitchen-Demos release.
Extraction rule
For every source episode:
Restore every recorded MuJoCo state with that episode's exact model.xml.gz
and ep_meta.json.
Find the first recorded contact between a Panda… See the full description on the dataset page: https://huggingface.co/datasets/Tsaochengyu/datasetSEGMENTfordraw.ReactiveGWM-Datasets
ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models
📚 Datasets-Introduction
ReactiveGWM-Datasets is the strategy-aligned training corpus that powers
ReactiveGWM, a game world
model that decouples player control from NPC autonomy. To learn that
decoupling, the model needs supervision that pairs each gameplay clip with
both a per-frame action stream (what the player did) and a high-level
NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.videlseal_eval_datasets
Merged Video Benchmark Captions and Semantic Embeddings
This dataset contains merged semantic caption segments and aligned text embeddings for four video benchmarks from a local DynamicvideoAgentRL/LVU cache. It includes captions and embeddings only, not source videos.
Files
captions.parquet: one row per caption segment.
semantic_vectors.float32.npy: NumPy array with shape 244909 x 3072; row i matches captions.parquet row where row_id == i.… See the full description on the dataset page: https://huggingface.co/datasets/CewEhao/videlseal_eval_datasets.datasetsWeisenV1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper",
"total_episodes": 77,
"total_frames": 17402,
"total_tasks": 5,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:77"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/malneyugnfl/datasetsWeisenV1.isaaclab_ur7e_lerobot_datasets_groot_v21
isaaclab_ur7e_lerobot_datasets_groot_v21
LeRobot v2.1 exports (per-episode parquet + mp4, meta/modality.json included) for NVIDIA Isaac GR00T N1.x training.
One sub-folder per task; each sub-folder is a complete, self-contained LeRobot v2.1 dataset.
sub-folder
task
episodes
frames
source (LeRobot v3.0)
pour_cup_clean/
Grasp the cup, lift it, and pour it out.
150
69,642
DexSteer/isaaclab-ur7e-pour_cup_clean
grasp_pan_clean/
Grasp the frying pan by its handle and… See the full description on the dataset page: https://huggingface.co/datasets/DexSteer/isaaclab_ur7e_lerobot_datasets_groot_v21.gello_datasets_abs_joint_abs_joint_30hzThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5",
"total_episodes": 400,
"total_frames": 173026,
"total_tasks": 7,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliy98/gello_datasets_abs_joint_abs_joint_30hz.liveness-filtering-datasetsdataset_simple_task_pick_and_place_with_distractorsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 823,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/dataset_simple_task_pick_and_place_with_distractors.so101-vla-datasets
SO-101 VLA Datasets
Vision-Language-Action(VLA) 학습용 데이터셋 모음. SO-101 (SO-ARM101) 팔로워 로봇으로
직접 수집한 tabletop manipulation 데모를 LeRobot v3.0 포맷으로 정리했습니다.
하나의 repo 안에 객체(object) → 태스크(task) → 변형(variant) 계층으로 하위폴더를 나눠 담습니다.
Robot: so_follower (SO-ARM101), 6-DoF
Format: LeRobot v3.0 — data/*.parquet + videos/*.mp4 + meta/
Modalities: RGB 카메라 2대 + 관절 상태/액션 + 자연어 태스크 지시문
License: Apache-2.0 (필요시 변경 가능)
📁 구조
so101-vla-datasets/
├── cube/pick-place/ # "orange… See the full description on the dataset page: https://huggingface.co/datasets/kangkb7701/so101-vla-datasets.MCWM_datasets
MCWM VPT 7.x Canonical Dataset
This dataset is a reproducible 5,460-episode subset of the OpenAI Video
PreTraining (VPT) 7.x contractor demonstrations, converted to the canonical
MCWM action and timestamp schema.
Contents
5,460 MP4 videos at 640x360
30,133,637 frames and 28,202,719 aligned canonical action ticks
418.45 hours of video
contractor-grouped splits: 4,767 train, 296 validation, 397 test
exact frame PTS, per-episode manifests, checksums, and an audit… See the full description on the dataset page: https://huggingface.co/datasets/dcloud347/MCWM_datasets.
