CoolFace
Modelpublic

Miical/pi05-libero10-task8-sft-10demos

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes40downloads
Model Card

PI0.5 LIBERO-10 task-8 initial policy

This checkpoint is the deliberately weak initial policy used in the ReCap study on LIBERO-10 task 8, “put both moka pots on the stove.” It starts from `Miical/pi05-base` and is fully fine-tuned on 10 official task demonstrations for 20 epochs (300 optimizer steps).

Training configuration

SettingValue
Training episodes10
Training frames4,033
Epochs / optimizer steps20 / 300
Model GPUs8
Global / micro batch size256 / 16
Learning rate1e-4
Weight decay1e-5
Warmup5%

On the complete 50-episode task-8 evaluation, this policy succeeds on 8/50 episodes (16%). The intentionally moderate success rate leaves room to measure post-training improvement. Three ReCap iterations raise the best measured success rate to 23/50 (46%).

The associated cumulative dataset is `Miical/pi05-libero10-task8-recap`. Its first 10 episodes are the demonstrations used for this checkpoint.