Aikwed/lingbot-vla-pick-and-place-toy-to-bucket
016
LingBot-VLA: Pick and Place Toy to Bucket
This checkpoint is a full-parameter fine-tune of LingBot-VLA-4B on the `AlvinAi/pick_and_place_toy_to_bucket` LeRobot dataset.
Task instruction
pick and place toy to bucketThe training pipeline formats it as <bos>pick and place toy to bucket\n before tokenization.
Dataset and robot layout
- 82 episodes, 27,962 frames, 30 FPS
- Cameras:
observation.images.frontandobservation.images.side - State/action dimensions:
shoulder_pan.posshoulder_lift.poselbow_flex.poswrist_flex.poswrist_roll.posgripper.pos- Arm joints are trained as delta actions.
- The gripper is trained as an absolute action.
- The action horizon is 50 steps, with episode-tail padding excluded from both normalization statistics and the training loss.
The dataset's gripper values are approximately in the range 0–34. Hardware inference must use the same calibration and units; do not assume a 0–100 gripper range.
Fine-tuning setup
- Full-parameter fine-tuning of the VLM, vision encoder, and action expert
- 4 epochs
- Global batch size 256 on 4 GPUs
- Learning rate
2e-5, constant schedule - FP32/TF32 FSDP2 training
- Mean/std normalization from the accompanying
norm_stats.json
Included files
- Hugging Face model weights and tokenizer/processor configuration
lingbotvla_cli.yaml: resolved training configurationnorm_stats.json: task-specific normalization statisticsrobot_config.yaml: feature mapping used during training and inference
For inference, use norm_stats.json and robot_config.yaml from this model repository so that input padding, normalization, action restoration, and the absolute gripper output remain aligned with training.
