CoolFace
Modelpublic

duanxingjuan/openvla-7b-libero-spatial-lora-clean-10k

sourceHugging Facemitupdated 2mo agoView on Hugging Face
1likes8downloads
Model Card

OpenVLA 7B · LIBERO-Spatial clean LoRA 10k

This is a merged, full OpenVLA checkpoint fine-tuned on LIBERO-Spatial. It is not an adapter-only upload. The run was created to replace an under-trained 2k-step LoRA policy with a controlled, resumable 10k-step experiment and an honest closed-loop evaluation.

Result at a glance

ItemValue
Base model`openvla/openvla-7b`
Training datalibero_spatial_no_noops from `openvla/modified_libero_rlds`
AdaptationLoRA on all linear layers, then merged into the base model
Training10,000 optimizer steps, effective batch 16
Closed-loop evaluation31/50 successes (62.0%)
Evaluation protocolLIBERO-Spatial, 10 tasks × 5 episodes, seed 7, center_crop=True
Intended useResearch and simulation only

Demo rollout

[image]

Download the MP4 · Browse the 10-task recovery rollouts

The demo assets were regenerated from this exact published checkpoint after evaluation. They use the same official harness, suite, seed, and center-crop setting, with one episode per task. They are qualitative demonstrations; the reported 31/50 score comes from the earlier complete 5-episode-per-task evaluation.

Evaluation

Evaluation used OpenVLA's official experiments/robot/libero/run_libero_eval.py harness. Success is LIBERO's environment goal predicate, not a visual or manually assigned label. The run completed all 50 episodes with no caught episode exceptions.

#LIBERO-Spatial taskSuccesses
1Black bowl between the plate and ramekin → plate1/5 (20%)
2Black bowl next to the ramekin → plate5/5 (100%)
3Black bowl from table center → plate1/5 (20%)
4Black bowl on the cookie box → plate5/5 (100%)
5Black bowl in the top drawer of the wooden cabinet → plate2/5 (40%)
6Black bowl on the ramekin → plate4/5 (80%)
7Black bowl next to the cookie box → plate4/5 (80%)
8Black bowl on the stove → plate2/5 (40%)
9Black bowl next to the plate → plate4/5 (80%)
10Black bowl on the wooden cabinet → plate3/5 (60%)
Overall10 tasks × 5 episodes31/50 (62.0%)

The 62% result crossed the experiment's predefined 60% usable threshold, but not its 75% stretch target. It also remains below a separate 84.9% reproduction of the official LIBERO-Spatial checkpoint. Episode counts differ, so the comparison should be read as directional rather than as a confidence-adjusted leaderboard.

Training recipe

SettingValue
OpenVLA source commitc8f03f48af692657d3060c19588038c7220e9af9
LoRA target modulesall-linear
LoRA rank / alpha / dropout32 / 16 / 0.0
Trainable parameters110,828,288 / 7,652,065,472 (1.45%)
Per-device batch / accumulation8 / 2
Effective batch16
OptimizerAdamW
Learning rateconstant 5e-4
Image augmentationenabled for all 10,000 steps
Gradient clippingglobal norm 1.0
Checkpoint intervalevery 1,000 optimizer steps, including optimizer and RNG state
Training hardware1× NVIDIA A100-SXM4-80GB
Stable throughputapproximately 1.31 seconds/step after CPU quota tuning

The final 500-step window averaged 65.34% action-token accuracy, 1.244 cross-entropy loss, and 0.0362 L1 action loss. No NaN or infinite training metrics were observed. The step-10000 adapter was merged into the base model, and the checkpoint includes the dataset_statistics.json needed for action un-normalization.

Loading the checkpoint

OpenVLA predicts one 7-DoF action at a time. A robot or simulator must call the policy repeatedly in a closed-loop controller.

python
import torch
from PIL import Image
from transformers import AutoModelForVision2Seq, AutoProcessor

model_id = "duanxingjuan/openvla-7b-libero-spatial-lora-clean-10k"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForVision2Seq.from_pretrained(
    model_id,
    attn_implementation="flash_attention_2",
    torch_dtype=torch.bfloat16,
    low_cpu_mem_usage=True,
    trust_remote_code=True,
).to("cuda:0")

image = Image.open("observation.png").convert("RGB")
instruction = "pick up the black bowl next to the ramekin and place it on the plate"
prompt = f"In: What action should the robot take to {instruction.lower()}?\nOut:"
inputs = processor(prompt, image).to("cuda:0", dtype=torch.bfloat16)

action = model.predict_action(
    **inputs,
    unnorm_key="libero_spatial_no_noops",
    do_sample=False,
)
print(action)  # shape: (7,)

For faithful LIBERO evaluation, use the official harness with --center_crop True; the center crop matches the image augmentation used throughout fine-tuning.

Limitations and safety

  • —This model was adapted and evaluated only on the simulated LIBERO-Spatial suite.
  • —Per-task results vary from 20% to 100%; the aggregate score hides meaningful spatial failure modes.
  • —Five trials per task is sufficient for this experiment's gate but still gives wide per-task uncertainty.
  • —The checkpoint has not been validated on a physical robot, other cameras, other embodiments, or safety- critical control. Do not deploy it on real hardware without independent validation, constraints, and an emergency-stop system.
  • —The model inherits the general limitations of OpenVLA and its pretraining data.

Reproducibility artifacts

The accompanying study repository contains the locked recipe, clipped/resumable trainer, preflight checks, watchdog, strict evaluation-summary validator, upload verifier, and CPU-only tests: `nele-duan/openVLA-study`.

The archived evaluation summary records seed=7, center_crop=true, 50 completed episodes, and the ten per-task rates shown above. The original model upload was verified at commit 1766524cf02cafd329d2a91e5cccb3bd5b8a10f9 before the card and demonstration assets were added.

Citation

bibtex
@article{kim24openvla,
  title={OpenVLA: An Open-Source Vision-Language-Action Model},
  author={Moo Jin Kim and Karl Pertsch and Siddharth Karamcheti and Ted Xiao and Ashwin Balakrishna and Suraj Nair and Rafael Rafailov and Ethan Foster and Grace Lam and Pannag Sanketi and Quan Vuong and Thomas Kollar and Benjamin Burchfiel and Russ Tedrake and Dorsa Sadigh and Sergey Levine and Percy Liang and Chelsea Finn},
  journal={arXiv preprint arXiv:2406.09246},
  year={2024}
}

This is an independent learning experiment, not an official OpenVLA release.