CoolFace
Modelpublic

StarVLA/StarVLA-Qwen3vl4b-PIv3-RoboDojo

sourceHugging Faceapache-2.0updated 23d agoView on Hugging Face
0likes63downloads
Model Card

StarVLA QwenPI_v3 Qwen3-VL-4B for RoboDojo (100k)

This directory contains one StarVLA QwenPI_v3 checkpoint initialized from Qwen3-VL-4B-Instruct and trained on the 35-task RoboDojo LeRobot v2.1 mixture. The VLM, VLM interface, and action model were trained end to end; this is not a LoRA or adapter-only checkpoint.

Model details

ItemValue
FrameworkStarVLA QwenPI_v3
Base VLMQwen3-VL-4B-Instruct
Action model36-layer LayerwiseFM
Action representation14D absolute joint position (abs_qpos)
Action horizon50
State dimension14
Camera inputHead, left wrist, right wrist; resized to 224 x 224
Inference flow steps4
Checkpoint step100,000
Checkpoint formatComplete StarVLA framework state dict (.pt)

The policy predicts a normalized 50 x 14 action chunk. RoboDojo evaluation must use the saved arx_x5 normalization statistics and execute 16 actions before requesting the next chunk.

Files

text
README.md
config.yaml
config.full.yaml
dataset_statistics.json
summary.jsonl
checkpoints/
└── steps_100000_pytorch_model.pt

Only the requested 100k checkpoint is included. Keep the configuration and dataset statistics beside the checkpoints/ directory; StarVLA uses them to reconstruct the framework and unnormalize actions.

Training details

SettingValue
Dataset mixturerobodojo_v21_all_h50_q99
Training tasks35
Per-GPU batch size16
Gradient accumulation1
Frozen modulesNone
OptimizerAdamW, betas (0.9, 0.95), epsilon 1e-8
VLM learning rate1e-5
VLM-interface learning rate1e-5
Action-model learning rate1e-4
ScheduleCosine, 5,000 warmup steps, minimum LR 5e-7
Gradient checkpointingEnabled
Random seed42

Official RoboDojo evaluation

All policies below use the official complete 42-task protocol: 50 episodes per task, 2,100 episodes per policy. Values are shown as SR (%) / Score. This directory's policy is bolded. Higher is better for both SR and Score.

Group summary

PolicyAverageGeneralizationPrecisionLong-HorizonMemoryOpen
QwenOFT4.86 / 8.014.33 / 6.4211.75 / 17.545.50 / 12.951.67 / 1.770.50 / 0.60
QwenGR00T3.81 / 7.353.50 / 6.525.75 / 10.096.50 / 15.463.33 / 4.370.00 / 0.00
QwenPI_v36.19 / 9.604.17 / 7.2814.00 / 19.0610.00 / 17.842.00 / 2.320.75 / 0.88

Task details

Task values are SR (%) / Score; each task uses 50 episodes. The bolded column is the policy in this directory.

Evaluation group / taskQwenOFTQwenGR00T**QwenPI_v3**
Generalization4.33 / 6.423.50 / 6.524.17 / 7.28
stack_bowls18.00 / 21.0010.00 / 14.8014.00 / 16.70
push_T0.00 / 0.000.00 / 0.000.00 / 0.00
packobjectsinto_box0.00 / 3.100.00 / 7.800.00 / 6.80
fold_clothes10.00 / 12.808.00 / 12.402.00 / 9.60
hang_mugs0.00 / 3.600.00 / 3.000.00 / 3.50
sweep_blocks0.00 / 0.000.00 / 0.002.00 / 2.00
pourliquidinto_cup14.00 / 14.0014.00 / 14.0012.00 / 12.00
make_toast0.00 / 1.002.00 / 5.002.00 / 5.00
arrangelargestnumber0.00 / 1.902.00 / 4.102.00 / 5.70
sortnestingdollsbysize0.00 / 0.004.00 / 4.006.00 / 6.00
storelaptopand_headphones4.00 / 11.200.00 / 7.202.00 / 8.40
stack_blocks6.00 / 8.402.00 / 5.908.00 / 11.60
Precision11.75 / 17.545.75 / 10.0914.00 / 19.06
fasten_screws4.00 / 8.000.00 / 2.000.00 / 6.00
plugincharger6.00 / 6.002.00 / 2.004.00 / 4.00
insert_tubes40.00 / 51.6028.00 / 40.4044.00 / 56.80
pourballsinto_vase8.00 / 8.000.00 / 0.002.00 / 2.00
play_Xylophone0.00 / 0.000.00 / 0.000.00 / 0.00
deposit_coin0.00 / 3.202.00 / 5.606.00 / 7.60
insert_key0.00 / 12.900.00 / 9.900.00 / 11.10
build_tower36.00 / 50.6014.00 / 20.8056.00 / 65.00
Long-Horizon5.50 / 12.956.50 / 15.4610.00 / 17.84
putbottlesinto_dustbin22.00 / 40.9026.00 / 44.4064.00 / 73.60
fillpenholder4.00 / 11.706.00 / 14.408.00 / 23.00
classify_objects2.00 / 5.506.00 / 11.500.00 / 7.50
playtictac_toe0.00 / 12.402.00 / 16.400.00 / 6.80
filleggholder0.00 / 0.600.00 / 0.000.00 / 0.80
organize_table0.00 / 16.500.00 / 25.004.00 / 27.00
make_kong16.00 / 16.0012.00 / 12.004.00 / 4.00
playstackingtoy0.00 / 0.000.00 / 0.000.00 / 0.00
Memory1.67 / 1.773.33 / 4.372.00 / 2.32
cover_blocks0.00 / 0.606.00 / 12.100.00 / 1.50
matchandpickfromconveyor10.00 / 10.0014.00 / 14.0012.00 / 12.00
swap_blocks0.00 / 0.000.00 / 0.000.00 / 0.00
swap_T0.00 / 0.000.00 / 0.000.00 / 0.00
pressbynumber0.00 / 0.000.00 / 0.000.00 / 0.00
imitatesortingsequence0.00 / 0.000.00 / 0.100.00 / 0.40
Open0.50 / 0.600.00 / 0.000.75 / 0.88
align_blocks0.00 / 0.000.00 / 0.000.00 / 0.00
general_pickup4.00 / 4.000.00 / 0.006.00 / 6.00
stackblocksby_language0.00 / 0.800.00 / 0.000.00 / 0.80
solve_equation0.00 / 0.000.00 / 0.000.00 / 0.00
classifyobjectsby_language0.00 / 0.000.00 / 0.000.00 / 0.20
pickfromconveyorbyimage0.00 / 0.000.00 / 0.000.00 / 0.00
storetoolsin_toolbox0.00 / 0.000.00 / 0.000.00 / 0.00
pourbylanguage0.00 / 0.000.00 / 0.000.00 / 0.00

Evaluation

This is a StarVLA checkpoint, not a Hugging Face from_pretrained() directory. Use the StarVLA model server and the XPolicyLab StarVLA adapter.

Start the model server from the StarVLA repository:

bash
huggingface-cli download StarVLA/StarVLA-Qwen3vl4b-PIv3-RoboDojo \
  --local-dir StarVLA-Qwen3vl4b-PIv3-RoboDojo
export CKPT="$PWD/StarVLA-Qwen3vl4b-PIv3-RoboDojo/checkpoints/steps_100000_pytorch_model.pt"
python deployment/model_server/server_policy.py \
  --ckpt_path "$CKPT" \
  --port 57700 \
  --use_bf16

Run a RoboDojo task from XPolicyLab/policy/starVLA:

bash
STARVLA_CKPT_PATH="$CKPT" \
STARVLA_INCLUDE_STATE=True \
STARVLA_UNNORM_KEY=arx_x5 \
STARVLA_EXECUTE_HORIZON=16 \
bash eval.sh \
  RoboDojo build_tower qwenpi_v3_steps_100000 \
  arx_x5 joint 0 0 1 <policy_conda_env> <robodojo_conda_env>

The final arguments are the seed, policy GPU, simulator GPU, policy environment, and RoboDojo environment. Use the RoboDojo task registry's native episode counts when producing an official aggregate.

Evidence and evaluation boundary

  • —Architecture, input/action dimensions, normalization contract, training settings, and the 100k release were checked against config.full.yaml, dataset_statistics.json, and the Hub file tree.
  • —The official-protocol table is retained from this published reference Card and matches the QwenPI_v3 100k checkpoint. This repository does not include raw per-episode evaluation logs, so the aggregate is not independently reconstructable from the repository files alone.
  • —Rows for QwenOFT and QwenGR00T are comparison context; they are not results produced by the checkpoint in this repository.
  • —Evaluation requires the same arx_x5 statistics, state inclusion, camera order, absolute-joint action mapping, and 16-step execution horizon.

Intended use and limitations

This checkpoint is intended for RoboDojo simulation research with the ARX X5 dual-arm embodiment. Performance with different camera calibration, state/action ordering, normalization, robots, or real-world hardware has not been established.

<!-- End of document -->

License evidence

The Apache-2.0 metadata above is retained from this target repository's previously published Model Card; it is not inferred from the StarVLA code license. No separate LICENSE file is packaged, and base-model and dataset terms remain applicable.