CoolFace
Modelpublic

Dreamer-VLA/dreamer-vla-lapa-aux-libero

sourceHugging Facecc-by-nc-4.0updated 5mo agoView on Hugging Face
0likes6downloads
Model Card

LAPA + inverse-dynamics aux finetune on LIBERO

LAPA's LAQ encoder (laq_openx, 344M params) unfrozen and finetuned on LIBERO with the inverse-dynamics auxiliary loss as the only supervision. 30k steps, batch 2, single GPU. Parallels the V-JEPA+aux pipeline for the LAPA family in the dreamer-vla paper.

Results (LIBERO task-OOD action probe, H=1, mean of 3 seeds)

Variantparamstrain R²test R²
LAPA frozen (reference)344M+0.240+0.410
LAPA + aux finetune ⭐344M+0.489+0.510

aux loss adds +0.10 over frozen baseline — comparable to V-JEPA 2 ViT-L frozen → +aux jump (+0.40 → +0.85, +0.44) but smaller magnitude.

Training recipe

  • —Init from laq_openx.pt (LAPA's pretrained LAQ on Open X-Embodiment)
  • —Encoder unfrozen (requires_grad=True)
  • —aux head: InvDynAuxHead(in_dim=1024, action_dim=7, hidden=[512,256])
  • —Loss: MSE between predicted action (from feature pair (ft, f{t+1})) and ground-truth LIBERO 7-D action
  • —AdamW lr=5e-5, weight_decay=0.05, cosine schedule with 2k warmup
  • —30k steps, batch 2 × clip_len 32 × LAPA's 256×256 image resolution
  • —Loss is the ONLY supervision (aux_lambda=1.0) — no separate SSL loss

Caveats

  • —LAPA's LAQ was trained on T=2 frame pairs; we feed T=32 by chunking into 31 adjacent pairs and averaging per-frame features. Slightly extrapolates beyond LAQ's training distribution but works well.
  • —344M params + batch=2 was the largest stable config on a single B200.
  • —Encoder finetune is destructive — this ckpt no longer matches LAPA's Open-X features. Use only as an aux-finetuned LIBERO encoder.

How to load

python
import sys, torch
sys.path.insert(0, "external_models/LAPA/laq")
from laq_model.latent_action_quantization import LatentActionQuantization

state = torch.load("ckpt_last.pt", weights_only=False)
model = LatentActionQuantization(
    dim=1024, quant_dim=32, codebook_size=8,
    image_size=256, patch_size=32,
    spatial_depth=8, temporal_depth=8,
    dim_head=64, heads=16, code_seq_len=4,
)
model.load_state_dict(state["encoder"])

The aux head is also in the ckpt (state["aux_head"]) but is training-only and ignored at probe time.

Citation

bibtex
@inproceedings{Ye2025LAPA,
  title={Latent Action Pretraining from Videos},
  author={Ye, Seonghyeon and others},
  year={2024}, booktitle={NeurIPS}
}

(dreamer-vla paper citation when on arXiv.)

Companion