CoolFace
Modelpublic

Aikwed/lingbot-vla-pick-and-place-toy-to-bucket

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes16downloads
Model Card

LingBot-VLA: Pick and Place Toy to Bucket

This checkpoint is a full-parameter fine-tune of LingBot-VLA-4B on the `AlvinAi/pick_and_place_toy_to_bucket` LeRobot dataset.

Task instruction

text
pick and place toy to bucket

The training pipeline formats it as <bos>pick and place toy to bucket\n before tokenization.

Dataset and robot layout

  • —82 episodes, 27,962 frames, 30 FPS
  • —Cameras: observation.images.front and observation.images.side
  • —State/action dimensions:
  • —shoulder_pan.pos
  • —shoulder_lift.pos
  • —elbow_flex.pos
  • —wrist_flex.pos
  • —wrist_roll.pos
  • —gripper.pos
  • —Arm joints are trained as delta actions.
  • —The gripper is trained as an absolute action.
  • —The action horizon is 50 steps, with episode-tail padding excluded from both normalization statistics and the training loss.

The dataset's gripper values are approximately in the range 0–34. Hardware inference must use the same calibration and units; do not assume a 0–100 gripper range.

Fine-tuning setup

  • —Full-parameter fine-tuning of the VLM, vision encoder, and action expert
  • —4 epochs
  • —Global batch size 256 on 4 GPUs
  • —Learning rate 2e-5, constant schedule
  • —FP32/TF32 FSDP2 training
  • —Mean/std normalization from the accompanying norm_stats.json

Included files

  • —Hugging Face model weights and tokenizer/processor configuration
  • —lingbotvla_cli.yaml: resolved training configuration
  • —norm_stats.json: task-specific normalization statistics
  • —robot_config.yaml: feature mapping used during training and inference

For inference, use norm_stats.json and robot_config.yaml from this model repository so that input padding, normalization, action restoration, and the absolute gripper output remain aligned with training.