CoolFace
Modelpublic

pengyue-polaron/lingbot-va-libero-goal-step-4000

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes
Model Card

LingBot-VA LIBERO-Goal — Step 4000

LingBot-VA fine-tuned on LIBERO-Goal for 4,000 steps. This release contains the transformer weights and configuration.

Evaluation

The checkpoint obtained 96.4% success (482/500) on LIBERO-Goal: 50 rollouts for each of the 10 tasks. Evaluation used 20 video denoising steps and 50 action denoising steps.

Task IDSuccessesEpisodesSuccess rate
05050100%
15050100%
2385076%
3495098%
4495098%
55050100%
65050100%
75050100%
85050100%
9465092%

Training

433 demonstrations, two 128×128 camera views, and 4,000 updates on 4 × H200. Full configuration and training recipe.

Epoch estimate (provisional)

If all 433 demonstrations each contribute exactly one valid training segment, the sample-exposure estimate is approximately 1,108.55 equivalent fine-tuning epochs: 4,000 optimizer updates × 120 effective global batch size / 433. The run-specific valid-segment list is not included, so this number remains conditional. Distributed-sampler padding can also make the actual DataLoader pass count differ. This estimate excludes base-model pretraining; see TRAINING.md for the reported batch configuration.

Loading

Download this repository and replace the base model's transformer/ directory with the transformer/ directory from this checkpoint while retaining the base model's tokenizer/, text_encoder/, and vae/.

Limitations

Evaluation covers LIBERO-Goal only. This is a multi-step LingBot-VA model, not a one-step Flash-WAM distillation.