CoolFace
Modelpublic

s1ghhh/VLADrop_pi05_LIBERO_Drop16ActionBlock_Endpoints

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes
Model Card

VLADrop-pi05-LIBERO-keep2-action

Checkpoint for Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?.

DTR (Drop-Then-Recovery) removes transformer blocks from a pretrained VLA model and recovery-fine-tunes the smaller dense model. Code: https://github.com/s1ghhh/VLADrop

This checkpoint

Paper rowTable 1: pi0.5 Keep 2 Action
Dropped blocksAction expert (Gemma, 18 layers): keep only blocks [0,17] (first & last), drop all others. Vision and language untouched.
Recovery trainingbatch size 32, 30K steps, lr 5e-5
LIBERO success rateSpatial 44.4 / Object 16.0 / Goal 40.8 / Long 3.6 / Avg 26.2 (per-suite values from evaluation logs)

Usage

This is an openpi-format pi0.5 checkpoint (PyTorch). Use with the VLADrop fork: https://github.com/s1ghhh/VLADrop

bash
python scripts/serve_policy_batch_drop.py \
    --config pi05_libero_dropped \
    --dir <this_repo_local_path> \
    --port 8000

Important: the drop lists are NOT stored inside the checkpoint. Pass the exact llm_drop_attn_list / llm_drop_mlp_list shown above (via config or CLI) when serving, otherwise layers will be mismatched. assets/ contains the LIBERO norm stats. The optimizer state (train_state/) is not included.

Citation

bibtex
@article{sun2026vladrop,
  title={Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?},
  author={Sun, Guoheng and Feng, Kaixi and He, Shwai and Gong, Xiaochuan and He, Yexiao and Wang, Ziyao and Shen, Zheyu and Ye, Wanghao and Kompella, Ramana Rao and Liu, Gaowen and Li, Ang},
  journal={arXiv preprint arXiv:2606.27755},
  year={2026}
}