mengqz9/deva_yam
DeVA — YAM (processed) Real bimanual-robot data collected on the YAM platform and processed for DeVA, a robot video-world model with affordance + depth physical guidance. Episodes 90 Frames 30,749 Video 480 × 1920, 30 fps (3 views tiled side by side) Views left_wrist, head, right_wrist Action dim 14 (bimanual, delta actions at train time) Proprio state included (state.json) Size 19 GB Download hf download mengqz9/deva_yam --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mengqz9/deva_yam.
DeVA — YAM (processed)
Real bimanual-robot data collected on the YAM platform and processed for DeVA, a robot video-world model with affordance + depth physical guidance.
Download
hf download mengqz9/deva_yam --repo-type dataset --local-dir datasets/deva_yamLayout
deva_yam/
├── dataset_info.json # view names, tile grid, affordance/depth encodings
├── videos/ # <episode>.mp4 tiled multi-view RGB
├── metas/ # <episode>.txt language instruction
├── t5_xxl/ # <episode>.pickle precomputed T5-XXL embedding [n_tok, 1024]
├── action.json # {episode: [[a_0..a_13], ...]} len == video frames
├── state.json # {episode: [[s_0..s_13], ...]} proprioceptive state
├── norm_stats.json # action normalization statistics (quantile)
├── affordance/ # <episode>.npz physical guidance (optional at train time)
└── depth/ # <episode>.npz physical guidance (optional at train time)videos/ tiles the views on the grid given by dataset_info.json::tile:
[[left_wrist, head, right_wrist]]The gripper channels of action.json are binarized (open/closed), which is intentional for this dataset.
Affordance / depth .npz
Both are keyed by canonical view name (never by tile position), one array per view:
T equals the episode's video frame count, so aux arrays index by absolute frame. The validity mask is derived from the tile grid (null cells are invalid) and is not stored.
Depth source — DepthAnything (ViT-L) disparity, converted as depth = 1 / disparity; higher value = farther. Clipped at export (percentile not recorded); ~5% of pixels saturate at 1.0.
Affordance source — UAD pseudo-labels rendered as Gaussian heatmaps. Note that the real-robot affordance footprint is noticeably broader than the simulator oracle labels used for LIBERO / RoboCasa.
Citation
If you use this data, please cite DeVA:
@article{zhang2026deva,
title = {{DeVA}: Decoupled Video-Action Model with physical guidance for robot policy learning},
author = {Zhang, Mengqi and Khose, Sahil and Kareer, Simar and
Song, Yuchen and Jain, Unnat and Hoffman, Judy},
journal = {arXiv preprint arXiv:2607.24159},
year = {2026}
}