arrow-hf/smolvla-robotwin-place-container-plate-50ep-multi
SmolVLA RoboTwin place_container_plate (50 ep, MULTI-instruction)
SmolVLA policy fine-tuned on 50 demonstration episodes of the `place_container_plate` task from RoboTwin 2.0 (demo_clean config), with per-episode random language instructions sampled from RoboTwin's 100 instruction variations (seed=42 for reproducibility).
This is the multi-instruction counterpart to `arrow-hf/smolvla-robotwin-place-container-plate-50ep` (which uses a single fixed instruction).
Task
- Robot: Agilex dual-arm, end-effector control (16D state, 16D action)
- Cameras: 3 RGB streams —
dual_cam_global,cam_wrist_65,cam_wrist_75(240×320, D435) - Control rate: ~30 Hz (LeRobot metadata is 10 Hz; underlying RoboTwin sim ~30 Hz, used consistently for train/eval)
- Instructions: 50 unique sentences (one per episode), examples:
- "Use the left arm to place the object into the basket"
- "Pick the item up and drop it into the woven basket"
- "Move the object from the table into the basket"
Training
Evaluation: Single vs Multi-Instruction Comparison
Evaluated in RoboTwin 2.0 simulator (demo_clean config), 10 episodes, max_steps=400, action_chunk_exec=50, single fixed eval instruction "place the container on the plate" (fair comparison).
The multi-instruction model trades some single-instruction performance for the ability to follow varied language commands. For tasks where instruction diversity helps (held-out new instructions), this trade-off may pay off.
Usage
from lerobot.policies.smolvla import SmolVLAPolicy
policy = SmolVLAPolicy.from_pretrained("arrow-hf/smolvla-robotwin-place-container-plate-50ep-multi")See LeRobot documentation for inference setup.
Citation
Built on SmolVLA and SmolVLA-RoboTwin pretrained base, fine-tuned on data collected from RoboTwin 2.0.
