makermods/smolvla_3cam_dagger_corr20_blue_cube_orange_tray
smolvla3camdaggercorr20bluecubeorange_tray
DAgger correction pass over `makermods/smolvla_3cam_200ep_blue_cube_orange_tray` checkpoint 15000, fine-tuned on 20 human-corrected rollout episodes. SO-101 (6-DoF), 3 cameras: "pick up blue cube and place in orange tray".
Cameras
observation.images.front / .wrist / .top, all 480x640. No --rename_map needed at inference — feed the real names directly.
Training
Caveats
91 epochs over 1.9 minutes of data. Prior correction passes in this project ran 28.6 and 48.8 epochs on larger sets. Loss was already flat by ~step 3000 (0.121 -> 0.110 over the final 2,000 steps), so the later checkpoints buy little and may overfit the corrections at the expense of the clean demos. Evaluate the earlier checkpoints — checkpoints/ holds one every 500 steps for exactly this reason.
The correction set's task string was recorded as "Pick up blue cube and place in orange box", which does not match the base model's training instruction. It was normalised to "pick up blue cube and place in orange tray" before training; the source dataset is unchanged.
