CoolFace
Modelpublic

Beeface/smolvla-libero-spatial

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes102downloads
Model Card

SmolVLA Fine-tuned on LIBERO-Spatial

This is a fine-tuned version of lerobot/smolvla_base trained on the LIBERO-Spatial benchmark using the LeRobot framework.

Demo Video

Task 8 success episode (70% success rate on this task):

<video controls> <source src="https://huggingface.co/Beeface/smolvla-libero-spatial/resolve/main/videos/task8success.mp4" type="video/mp4"> </video>

Model Details

  • —Base model: lerobot/smolvla_base
  • —Parameters: 450M total (100M trainable action expert)
  • —Training steps: 20,000
  • —Batch size: 8
  • —Hardware: NVIDIA L4 24GB (Google Colab Pro)
  • —Training time: ~2.5 hours

Performance on LIBERO-Spatial

TaskSuccess Rate
task_060%
task_150%
task_260%
task_310%
task_420%
task_520%
task_610%
task_730%
task_870%
task_930%
Overall36%

Training Command

bash
lerobot-train \
  --policy.type=smolvla \
  --policy.pretrained_path=lerobot/smolvla_base \
  --dataset.repo_id=HuggingFaceVLA/libero \
  --batch_size=8 \
  --steps=20000 \
  --seed=42

Ablation Study — Training Duration

We evaluated checkpoints at multiple steps to understand convergence:

Training StepsSuccess Rate
2,0002%
6,00017%
10,00031%
20,00036%

Performance improves consistently but with diminishing returns, suggesting convergence begins around 10K steps on LIBERO-Spatial.

Framework