khanhnd61/octo-small_so101-multi-task-clean
Model Card for octo
<!-- Provide a quick summary of what the model is/does. -->
This is a octo policy trained with LeRobot.
<!-- A short demo is worth more than any description! Record a GIF/video of the policy running on your robot, upload it to this repo, and embed it here: <p align="center"> <img src="https://huggingface.co/<hfuser>/<policyrepo_id>/resolve/main/demo.gif" width="60%"/> </p> -->
This policy has been trained and pushed to the Hub using LeRobot.
See the full LeRobot documentation.
What changed in this revision
Retrained at Octo-Small's documented recipe (batch_size=32, steps=20000, per src/lerobot/policies/octo/README.md) with image augmentation enabled, replacing a previous batch_size=16 / 9500-step run that was an ACT-style budget and left this model underfit (9.9 epochs against the documented 41.8).
Measured on 6 held-out episodes (2 per task, eval_split=0.1, never seen in training):
The jerk figure is the one that matters here. Octo re-plans every 4 frames (7.5 Hz at 30 FPS) and its diffusion head resamples each time, so sampling variance reaches the servos as visible shaking. This revision is 3.5x smoother while being slightly more accurate, and is smoother than an ACT-based policy trained on the same data (0.455).
Two notes for anyone reproducing this:
- *The last checkpoint is the right one, and held-out loss says otherwise.* Diffusion eval loss was lowest around step 6000 (0.538) and worst at step 20000 (0.952), yet step 20000 is dramatically the best on actual sampled actions - step 4000 produces a jerk of 7.29. Average denoising error is a poor proxy for sample quality here; score the actions.
- `n_inference_samples` defaults to 8 in this checkpoint. Octo's reference implementation draws one sample per chunk; averaging 8 cuts sampling jitter about 3x and costs almost nothing, because the transformer runs once and only the small score network repeats (13.7 -> 13.5 ms per re-plan). Override with
--policy.n_inference_samples=1to reproduce upstream exactly.
Known data defect
The follower arm's wrist_roll never tracked the leader during recording: the action column spans 111-163 deg while the recorded state sits frozen at ~87.5 deg (per-episode std 0.00, action/state correlation about 0). Every policy trained on this dataset commands wrist_roll 60-70 deg beyond what the arm can reach, on 100% of timesteps. Inert for task behaviour, but it drives that servo into its stop for the whole rollout.
Model Details
- License: apache-2.0
- Robot type:
so_follower - Cameras:
front,wrist
Inputs & Outputs
The policy consumes these observation features and produces these action features.
Inputs
Outputs
Training Dataset
- Repository: khanhnd61/so101-multi-task-clean
- Episodes: 44
- Frames: 15317
- Frame rate: 30 FPS
- Task(s): "Put the tape into the box", "Put the tape into the cup", "Put the cup into the box"
<a class="flex" href="https://huggingface.co/spaces/lerobot/visualize_dataset?path=khanhnd61/so101-multi-task-clean"> <img class="block dark:hidden" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl.svg"/> <img class="hidden dark:block" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl-dark.svg"/> </a>
Training Configuration
How to Get Started with the Model
New to LeRobot? These guides cover the full workflow:
- [Install LeRobot](https://huggingface.co/docs/lerobot/main/en/installation) — set up the
lerobotpackage. - [Hardware setup](https://huggingface.co/docs/lerobot/main/en/hardware_guide) — assemble, wire, and calibrate your robot and cameras.
- [Record data & train a policy](https://huggingface.co/docs/lerobot/en/il_robots) — the end-to-end imitation-learning walkthrough.
- [CLI cheat-sheet](https://huggingface.co/docs/lerobot/main/en/cheat-sheet) — quick reference for the
lerobot-*commands.
The short version to run and train this policy:
Run the policy on your robot
lerobot-rollout \
--strategy.type=base \
--robot.type=so_follower \
--robot.port=<your_robot_port> \
--robot.cameras="{ <camera_1>: {type: opencv, index_or_path: <index_or_path>, width: 640, height: 480, fps: 30}, <camera_2>: {type: opencv, index_or_path: <index_or_path>, width: 640, height: 480, fps: 30}}" \
--policy.path=khanhnd61/octo-small_so101-multi-task-clean \
--task="Put the tape into the box" \
--duration=60Replace the remaining <...> placeholders with your own values: --robot.port and the camera names/indices are specific to your machine, and the camera names must match the observation keys this policy was trained on.
When --strategy.type=base is used the script doesn't record the episodes. Skipping duration will make the policy run indefinitely. For more information look at rollout documentation.
Train your own policy
lerobot-train \
--dataset.repo_id=${HF_USER}/<dataset> \
--policy.type=octo \
--output_dir=outputs/train/<policy_repo_id> \
--job_name=lerobot_training \
--policy.device=cuda \
--policy.repo_id=${HF_USER}/<policy_repo_id> \
--wandb.enable=trueWrites checkpoints to `outputs/train/<policyrepoid>/checkpoints/`.
Evaluation
<!-- Report real-robot results here: run the policy several times per task and count the successes. Delete the "No evaluation results" line and fill in this table instead:
Also worth noting: anything that affects difficulty (new object positions, lighting, distractors, a different robot of the same type, ...). -->
No evaluation results have been provided for this policy yet.
Citation
If you use this policy, please cite the method linked in the description above, along with LeRobot:
@misc{cadene2024lerobot,
author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and Gallouedec, Quentin and Zouitine, Adil and Palma, Steven and Kooijmans, Pepijn and Aractingi, Michel and Shukor, Mustafa and Aubakirova, Dana and Russi, Martino and Capuano, Francesco and Pascal, Caroline and Choghari, Jade and Moss, Jess and Wolf, Thomas},
title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch},
howpublished = "\url{https://github.com/huggingface/lerobot}",
year = {2024}
}