CoolFace
Modelpublic

io-intelligence/smolvla_so101_stack_cups

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes15downloads
Model Card

Model Card for smolvla

<!-- Provide a quick summary of what the model is/does. -->

SmolVLA is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.

Fine-tuned on SO-101 for the task "Stack the cups".

<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/640e21ef3c82bd463ee5a76d/aooU0a3DMtYmy_1IWMaIM.png" alt="smolvla architecture" width="85%"/> </p>

<p align="center"> <video src="https://huggingface.co/io-intelligence/smolvlaso101stackcups/resolve/main/inferencedemo.mp4" controls width="70%"></video> </p>

<p align="center"><em>Real-robot inference demo (also available as <code>inference_demo.mp4</code> in this repo).</em></p>

This policy has been trained and pushed to the Hub using LeRobot.

Learn how to train and run it in the LeRobot smolvla guide, or browse the full documentation.


Model Details

  • License: apache-2.0
  • Fine-tuned from: lerobot/smolvla_base
  • Robot type: so101_follower (SO-101)
  • Cameras (physical → policy): see table below

Camera mapping

The dataset records views as front / top / wrist. During SmolVLA training they are renamed to camera1 / camera2 / camera3. At inference you should keep the physical names on the robot and pass the same rename_map.

Physical cameraMount / rolePolicy feature after rename
frontFront-facing view of the workspace (Astra)observation.images.camera1
topTop-down / overhead view (Astra)observation.images.camera2
wristWrist / gripper cameraobservation.images.camera3
json
{
  "observation.images.front": "observation.images.camera1",
  "observation.images.top": "observation.images.camera2",
  "observation.images.wrist": "observation.images.camera3"
}

Inputs & Outputs

The policy consumes these observation features and produces these action features.

Inputs

FeaturePhysical sourceTypeShape
observation.stateJoint stateSTATE(6,)
observation.images.camera1frontVISUAL(3, 256, 256)
observation.images.camera2topVISUAL(3, 256, 256)
observation.images.camera3wristVISUAL(3, 256, 256)

Outputs

FeatureTypeShape
actionACTION(6,)

Training Dataset

<a class="flex" href="https://huggingface.co/spaces/lerobot/visualizedataset?path=io-intelligence/so101stack_cups"> <img class="block dark:hidden" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl.svg"/> <img class="hidden dark:block" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl-dark.svg"/> </a>

Training Configuration

SettingValue
Training steps30000
Batch size8
Optimizeradamw
Learning rate0.0001
Seed1000
LeRobot version0.6.1

How to Get Started with the Model

New to LeRobot? These guides cover the full workflow:

  • [Install LeRobot](https://huggingface.co/docs/lerobot/main/en/installation) — set up the lerobot package.
  • [Hardware setup](https://huggingface.co/docs/lerobot/main/en/hardware_guide) — assemble, wire, and calibrate your robot and cameras.
  • [Record data & train a policy](https://huggingface.co/docs/lerobot/en/il_robots) — the end-to-end imitation-learning walkthrough.
  • [CLI cheat-sheet](https://huggingface.co/docs/lerobot/main/en/cheat-sheet) — quick reference for the lerobot-* commands.

Run the policy on your robot

Use physical camera keys front / top / wrist, then apply the rename map so they match the policy's camera1 / camera2 / camera3 features. SmolVLA works best with RTC inference.

bash
lerobot-rollout \
  --strategy.type=base \
  --robot.type=so101_follower \
  --robot.port=<your_robot_port> \
  --robot.cameras="{ \
    front: {type: opencv, index_or_path: <front_device>, width: 640, height: 480, fps: 30}, \
    top: {type: opencv, index_or_path: <top_device>, width: 640, height: 480, fps: 30}, \
    wrist: {type: opencv, index_or_path: <wrist_device>, width: 640, height: 480, fps: 30} \
  }" \
  --policy.path=io-intelligence/smolvla_so101_stack_cups \
  --rename_map='{"observation.images.front":"observation.images.camera1","observation.images.top":"observation.images.camera2","observation.images.wrist":"observation.images.camera3"}' \
  --inference.type=rtc \
  --task="Stack the cups" \
  --duration=60

Replace <your_robot_port> and the three camera device paths with your machine values. Camera names must stay front / top / wrist (not camera1/2/3 on the robot side).

When --strategy.type=base is used the script doesn't record episodes. Set --duration=0 (or omit duration depending on your CLI) to run until Ctrl+C. For more information see the rollout / inference docs.

Train your own policy

This policy type is usually fine-tuned from the pretrained base model lerobot/smolvla_base:

bash
lerobot-train \
  --dataset.repo_id=${HF_USER}/<dataset> \
  --policy.path=lerobot/smolvla_base \
  --output_dir=outputs/train/<policy_repo_id> \
  --job_name=lerobot_training \
  --policy.device=cuda \
  --policy.repo_id=${HF_USER}/<policy_repo_id> \
  --wandb.enable=true

Writes checkpoints to `outputs/train/<policyrepoid>/checkpoints/`.


Evaluation

Real-robot inference demo of stacking cups is included as `inference_demo.mp4` (copied from the training dataset root).

TaskNotes
Stack the cupsQualitative success demo on SO-101 (see video above)

Citation

If you use this policy, please cite the method linked in the description above, along with LeRobot:

bibtex
@misc{cadene2024lerobot,
    author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and Gallouedec, Quentin and Zouitine, Adil and Palma, Steven and Kooijmans, Pepijn and Aractingi, Michel and Shukor, Mustafa and Aubakirova, Dana and Russi, Martino and Capuano, Francesco and Pascal, Caroline and Choghari, Jade and Moss, Jess and Wolf, Thomas},
    title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch},
    howpublished = "\url{https://github.com/huggingface/lerobot}",
    year = {2024}
}