StarVLA/VLAct_Qwen3OFT_Robotwin_all_Finetune
<p align="center"> <img src="https://raw.githubusercontent.com/starVLA/VLAct/main/assets/VLAct-update.png" width="92%" alt="VLAct overview: representation-centric continued pre-training for vision-language-action models"> </p>
VLAct · Qwen3-VL-4B OFT · RoboTwin 2.0 (All / Data Scaling)
    
This repository contains the 100K-step VLAct downstream fine-tuning checkpoint for RoboTwin 2.0 All (Data Scaling), introduced in Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models. It starts from `StarVLA/VLAct_Qwen3_Pretrain` and adapts the shared VLAct backbone with a randomly initialized OFT action head for the target benchmark.
[!IMPORTANT] This is a StarVLA training / evaluation checkpoint, not a standardtransformers.AutoModelpackage. Load it with the matching StarVLA framework (QwenOFT) and the packagedconfig.yaml/dataset_statistics.json. Safe robot deployment still requires embodiment-specific action mapping, normalization, camera calibration, control-rate handling, workspace constraints, and independent safety systems.
What is VLAct?
VLAct is a representation-centric continued-pretraining recipe for vision-language-action models. It preserves the VLM prior, co-trains multiple continuous action heads on a shared latent, and shares action semantics across embodiments with a partially unified padded layout and wrap-aware joint loss. Downstream policies discard the pretraining heads, randomly initialize a target action head, and fine-tune from the VLAct backbone.
The complete method, ablations, and evaluation protocols are documented in the paper and code repository.
Checkpoint details
Fine-tuning data
Fine-tuning uses the larger RoboTwin 2.0 All mixture (robotwin_all_wrap_32) with balance_dataset_weights=True. Absolute-joint actions, CoT object-grounding prompts, and wrap-aware angular-joint loss follow the RoboTwin OFT launcher.
Recommended use: download and evaluate
1. Install StarVLA / VLAct
git clone https://github.com/starVLA/VLAct.git
cd VLAct
conda create -n vlact python=3.10 -y
conda activate vlact
# Install a CUDA-compatible PyTorch build first.
python -m pip install -r requirements.txt
python -m pip install flash-attn==2.7.4.post1 --no-build-isolation
python -m pip install -e .2. Download the checkpoint
Run from the VLAct repository root:
hf download StarVLA/VLAct_Qwen3OFT_Robotwin_all_Finetune \
--local-dir playground/Pretrained_models/VLAct-Qwen3VL4B-OFT-Robotwin-AllThe checkpoint path is then:
playground/Pretrained_models/VLAct-Qwen3VL4B-OFT-Robotwin-All/checkpoints/steps_100000_pytorch_model.ptKeep the downloaded directory structure unchanged. StarVLA resolves config.yaml and dataset_statistics.json from the run directory two levels above the checkpoint file.
Action un-normalization
The policy predicts actions normalized to roughly [-1, 1]. StarVLA maps them back to physical units using the q01 / q99 / mask statistics stored in dataset_statistics.json, and the unnorm_key selects which statistics block to use.
This run packages a single key, new_embodiment. StarVLA resolves it automatically when only one key is present, so you normally do not need to set anything. If your evaluation config exposes an unnorm_key field (for example DOMINO's examples/DOMINO/eval_files/deploy_policy.yml or the eval config generated by the RoboTwin launcher), set it to new_embodiment.
Changing the normalization statistics, camera ordering, state usage, action ordering, or execution horizon can materially change results.
3. Evaluate with the StarVLA / VLAct scripts
Follow the RoboTwin evaluation guide in `examples/Robotwin/README.md`.
Reproduce training with:
bash scripts/run_scripts/RoboTwin/train_robotwin_qwen3oft.shLoading the policy
Reconstruct the policy with the matching StarVLA framework (QwenOFT) and the packaged configuration. The .pt file contains model parameters only; it does not package optimizer or scheduler state.
This checkpoint is intended for evaluation or further fine-tuning on the same embodiment / action contract. Transferring it to a different robot, camera setup, or action space usually requires additional adaptation.
Files
VLAct-Qwen3VL4B-OFT-Robotwin-All/
├── README.md
├── config.yaml
├── training_config.original.yaml
├── dataset_statistics.json
├── summary.jsonl
└── checkpoints/
└── steps_100000_pytorch_model.ptCheckpoint SHA-256:
beb2c0ec39c3bc7173a5bd90eb59ef99854b73ca1aa3bf50fb6ffc353f477c2aBenchmark results
This checkpoint matches the head (OFT) and setting of the VLAct RoboTwin 2.0 Data Scaling result reported in the paper:
92.5% is the default VLAct-OFT Clean success in the Data Scaling setting, and 90.8% is the result under randomization.
Use the paper and RoboTwin evaluation scripts for the exact protocol.
Intended use and limitations
This checkpoint is intended for research on VLA representation transfer and benchmark evaluation on RoboTwin 2.0 All (Data Scaling).
- It has not been validated as a universal zero-shot policy across arbitrary robots.
- Safe deployment requires embodiment-specific action mapping, normalization, camera calibration, control-rate handling, workspace constraints, and independent safety systems.
- Performance depends on evaluation protocol, simulator / real-robot setup, observation configuration, and action execution settings.
- The model may inherit limitations and biases from its base VLM, the VLAct pretraining mixture, and the downstream fine-tuning data.
Citation
@article{yang2026beyond,
title={Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models},
author={Yang, Senqiao and Wang, Chengyao and Chen, Yuxin and Wang, Zixuan and Tang, Longxiang and Gui, Haokun and Ye, Jinhui and Lu, Changsheng and Wu, Xiaoyang and Zhu, Mingkang and others},
journal={arXiv preprint arXiv:2608.27550},
year={2026}
}License and acknowledgements
The checkpoint is released under the Apache License 2.0. The VLAct code repository is released separately under the MIT License. Users must also comply with the licenses and terms of the base model, the VLAct pretraining checkpoint, and the training / evaluation datasets.
VLAct builds on StarVLA, LeRobot, GR00T, and Qwen3-VL.
For questions, email yangsenqiao.ai@gmail.com or open an issue in the VLAct repository.
