CoolFace
Datasetpublic

JackieMM/aloha_real_agilex_folding_fabric-vjepa2-pool

aloha_real_agilex_folding_fabric with pooled V-JEPA2.1 targets This is a LeRobot v2.1 copy of SakikoTogawa/aloha_real_agilex_folding_fabric with observation.concept added. The 2304-dimensional target follows vjepa2-1-vitb384-spatialpool-t0-3concat: for every timestep and each camera in observation.images.cam_left_wrist, observation.images.cam_right_wrist, observation.images.cam_high, a 64-frame forward clip is resized to 384x384, encoded with vjepa2_1_vit_base_384, spatially… See the full description on the dataset page: https://huggingface.co/datasets/JackieMM/aloha_real_agilex_folding_fabric-vjepa2-pool.

sourceHugging Faceupdated 18d agoView on Hugging Face
0likes88downloads
Dataset Card

aloharealagilexfoldingfabric with pooled V-JEPA2.1 targets

This is a LeRobot v2.1 copy of SakikoTogawa/aloha_real_agilex_folding_fabric with observation.concept added. The 2304-dimensional target follows vjepa2-1-vitb384-spatialpool-t0-3concat: for every timestep and each camera in observation.images.cam_left_wrist, observation.images.cam_right_wrist, observation.images.cam_high, a 64-frame forward clip is resized to 384x384, encoded with vjepa2_1_vit_base_384, spatially mean-pooled at temporal token 0, and the three 768-dimensional camera vectors are concatenated. Tail clips repeat the last frame. These targets are intended only as training-time supervision.