Beeface/smolvla-libero-spatial
0102
SmolVLA Fine-tuned on LIBERO-Spatial
This is a fine-tuned version of lerobot/smolvla_base trained on the LIBERO-Spatial benchmark using the LeRobot framework.
Demo Video
Task 8 success episode (70% success rate on this task):
<video controls> <source src="https://huggingface.co/Beeface/smolvla-libero-spatial/resolve/main/videos/task8success.mp4" type="video/mp4"> </video>
Model Details
- Base model: lerobot/smolvla_base
- Parameters: 450M total (100M trainable action expert)
- Training steps: 20,000
- Batch size: 8
- Hardware: NVIDIA L4 24GB (Google Colab Pro)
- Training time: ~2.5 hours
Performance on LIBERO-Spatial
Training Command
lerobot-train \
--policy.type=smolvla \
--policy.pretrained_path=lerobot/smolvla_base \
--dataset.repo_id=HuggingFaceVLA/libero \
--batch_size=8 \
--steps=20000 \
--seed=42Ablation Study — Training Duration
We evaluated checkpoints at multiple steps to understand convergence:
Performance improves consistently but with diminishing returns, suggesting convergence begins around 10K steps on LIBERO-Spatial.
