CoolFace
Modelpublic

RyanL22/pi05-rby1-wujihand2-teleopv1-baseline-20k

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes22downloads
Model Card

pi0.5 baseline — RB-Y1 + Wujihand2, real teleop only (teleop v1), step 20k

LeRobot-native pi05 (v0.6.1) fine-tuned only on real teleoperation data of the RB-Y1 mobile manipulator with Wujihand2 hands: 209 episodes / 108,021 frames at 30 fps (ball 51 / bottle 51 / box 55 / doll 52). No synthetic / human-video data.

Checkpoint: step 20,000 (final) of a 20,000-step run (Kakao, 4x A100 80GB, batch 16 x 4 = 64).

54-dim state / action (differs from lerobot/pi05_base)

The recorded robot vector is 66-dim. Wheels (4), torso (6) and head (2) are constant in this data and were dropped; the policy sees and predicts the remaining 54 dims = joint_position[10:64], in this order:

slicedimsjoints
0:77rightarm0, rightarm1, rightarm2, rightarm3, rightarm4, rightarm5, rightarm6
7:2720right hand: thumbcmcflex, thumbcmcabd, thumbmcp, thumbip, indexfingermcpflex, indexfingermcpabd, indexfingerpip, indexfingerdip, middlefingermcpflex, middlefingermcpabd, middlefingerpip, middlefingerdip, ringfingermcpflex, ringfingermcpabd, ringfingerpip, ringfingerdip, pinkymcpflex, pinkymcpabd, pinkypip, pinkydip
27:347leftarm0, leftarm1, leftarm2, leftarm3, leftarm4, leftarm5, leftarm6
34:5420left hand: same 20 joints as the right hand, left_ prefix

pi0.5 pads actions to max_action_dim = 32, so the base checkpoint was widened first: action_in_proj [1024, 32] -> [1024, 54] and action_out_proj [32, 1024] -> [54, 1024] (pretrained columns kept, new ones initialised from the same std, new bias 0), max_state_dim = 54, and the tokenizer max_length raised 200 -> 320 because 54 discretised state numbers are part of the prompt. All of that is already in this repo's config.json / policy_preprocessor.json.

At rollout: feed observation.state as the 54 values above (robot order, index 10..63) and write the 54 predicted values back to the same joints; hold wheels / torso / head at the recording pose.

Training settings

value
vision encoder (SigLIP, 412.4M)fine-tuned (not frozen)
image augmentationphotometric + affine, one draw replayed across the stereo pair
cell shares∝ sqrt(frames), 4 cells
mirror augmentationoff
optimizerAdamW, peak lr 2.5e-5, cosine decay to 2.5e-6, warmup 1000
precisionbfloat16, gradient checkpointing
chunk50 actions @ 30 fps (1.67 s), n_obs_steps=1
normalization statsquantiles; no constant dims remain (gate: q99 - q01 >= 1e-3)

Inputs

  • —observation.images.base_0_rgb <- left ZED view, resized 1280x720 -> 512x288
  • —observation.images.left_wrist_0_rgb <- right ZED view, 512x288
  • —observation.state — 54 dims (above) action — 54 dims, absolute joint targets

Tasks: pick up the {ball | bottle | small box | doll} from table and put it in the white box.

Load

python
from lerobot.policies.pi05.modeling_pi05 import PI05Policy
policy = PI05Policy.from_pretrained("RyanL22/pi05-rby1-wujihand2-teleopv1-baseline-20k")