CoolFace
Modelpublic

snupilab/humanoidtoolbench-psi0-sim-3003

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes
Model Card

HumanoidToolBench Psi0: Simulation training, 3,003 segments

The validated final policy checkpoint is available in this repository.

This is a HumanoidToolBench training-result repository. It does not substitute an upstream pretrained policy for a HumanoidToolBench-trained checkpoint.

SettingValue
Training stageSimulation training, 3,003 segments
Target optimizer updates40000
Per-GPU batch / GPUs / global batch16 / 8 / 128
Gradient accumulation1
Conditions per global batch18
Dataset revision8b2cd31e107b64cb13f812ea217a63a20845c78a

Pinned training data.

The simulation pool contains 1,200 successful L1/L2 demonstrations and 1,803 extracted L0 prefixes, spanning 18 conditions. The 3,003 segments are not 3,003 independent demonstrations.

Use the model's native HumanoidToolBench adapter and model-specific dependencies. This repository does not claim compatibility with arbitrary Transformers or simulation loaders. No evaluation score is claimed by checkpoint publication.

Use the pinned Psi0 native code and HumanoidToolBench adapter, run directory run and final checkpoint step. Change cwd to exported run for datasetstatistics.json resolution. Before invoking Psi0Model.frompretrained, set psi.models.psi0.QWEN3VLVARIANT to the absolute bundled processor directory (Qwen/Qwen3-VL-2B-Instruct@89644892e4d85e24eaac8bacfd4f463576704203). This overrides the native unpinned processor/config identifier without loading upstream VLM weights. The trained safetensors strictly loads both VLM and actionheader; native token embedding tying is retained. Keep the native one-camera preprocessing, instruction case, state/action normalization, RTC and H30 configuration.

Training uses independent model optimizers and shared GPU execution through MPS. Publication is performed by a CPU uploader after final checkpoint validation.