CoolFace
Modelpublic

makermods/smolvla_3cam_dagger_corr20_blue_cube_orange_tray

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes12downloads
Model Card

smolvla3camdaggercorr20bluecubeorange_tray

DAgger correction pass over `makermods/smolvla_3cam_200ep_blue_cube_orange_tray` checkpoint 15000, fine-tuned on 20 human-corrected rollout episodes. SO-101 (6-DoF), 3 cameras: "pick up blue cube and place in orange tray".

Cameras

observation.images.front / .wrist / .top, all 480x640. No --rename_map needed at inference — feed the real names directly.

Training

base3-cam checkpoint 15000 (not 20000 — 15k evaluated better)
corrections20 episodes / 3,510 frames (1.9 min), 100% intervention frames
steps5,000 @ batch 64 = 91 epochs over the correction set
lr1e-5 (10x below from-base), warmup auto-scaled 1000 -> 250, decay -> 5,000
loss2.193 (step 10) -> 0.110 (step 5000), grad norm 13.0 -> 1.17
hardwareRTX 4090, bf16 AMP, 19.9 GB, 1h09m

Caveats

91 epochs over 1.9 minutes of data. Prior correction passes in this project ran 28.6 and 48.8 epochs on larger sets. Loss was already flat by ~step 3000 (0.121 -> 0.110 over the final 2,000 steps), so the later checkpoints buy little and may overfit the corrections at the expense of the clean demos. Evaluate the earlier checkpoints — checkpoints/ holds one every 500 steps for exactly this reason.

The correction set's task string was recorded as "Pick up blue cube and place in orange box", which does not match the base model's training instruction. It was normalised to "pick up blue cube and place in orange tray" before training; the source dataset is unchanged.