JackieMM/aloha_real_agilex_folding_fabric-vjepa2-pool
aloha_real_agilex_folding_fabric with pooled V-JEPA2.1 targets This is a LeRobot v2.1 copy of SakikoTogawa/aloha_real_agilex_folding_fabric with observation.concept added. The 2304-dimensional target follows vjepa2-1-vitb384-spatialpool-t0-3concat: for every timestep and each camera in observation.images.cam_left_wrist, observation.images.cam_right_wrist, observation.images.cam_high, a 64-frame forward clip is resized to 384x384, encoded with vjepa2_1_vit_base_384, spatially… See the full description on the dataset page: https://huggingface.co/datasets/JackieMM/aloha_real_agilex_folding_fabric-vjepa2-pool.
aloharealagilexfoldingfabric with pooled V-JEPA2.1 targets
This is a LeRobot v2.1 copy of SakikoTogawa/aloha_real_agilex_folding_fabric with observation.concept added. The 2304-dimensional target follows vjepa2-1-vitb384-spatialpool-t0-3concat: for every timestep and each camera in observation.images.cam_left_wrist, observation.images.cam_right_wrist, observation.images.cam_high, a 64-frame forward clip is resized to 384x384, encoded with vjepa2_1_vit_base_384, spatially mean-pooled at temporal token 0, and the three 768-dimensional camera vectors are concatenated. Tail clips repeat the last frame. These targets are intended only as training-time supervision.
