CoolFace
Modelpublic

Elvinky/pi05-piperx-1h-intervention-1to1-mask-8066s-seed1000-fp32

sourceHugging Faceupdated 17d agoView on Hugging Face
0likes28downloads
Model Card

Pi0.5 PiperX: D1 + own human-intervention repair

Inference checkpoint after 8,066 new fine-tuning steps, initialized from the original 1h demonstration model at step 4,193. This is not the full558/50k model or the older 616-segment intervention experiment.

Task: Insert the copper screw into the black sleeve.

Data and training

  • —Replay: the original seed-1000 D1 selection, 54 episodes / 107,322 frames (~59m37s).
  • —Intervention source revision: d77909aac7383c625225f70e74a9a609b1712adf of the linked D1 full-episode dataset.
  • —All 103,237 human-controlled frames (~57m21s) extracted into 557 continuous segments. Automatic frames excluded. All 8 segments shorter than 50 frames retained; minimum length 9 frames.
  • —Exact 1:1 sampling: each rank uses 4 demo + 4 intervention samples; 8 GPUs, global batch 64.
  • —2.5 intervention exposure epochs: ceil(2.5 * 103237 / 32) = 8066 optimizer steps. Not 2.5 combined-dataset epochs.
  • —Full-parameter FP32; AMP and model compilation off; gradient checkpointing on. AdamW betas (0.9, 0.95), epsilon 1e-8, weight decay 0.01, clip norm 1.0.
  • —Learning rate 2.5e-6 to 2.5e-7, warmup 200, seed 1000. New optimizer/scheduler for fine-tuning. Only final checkpoint saved.
  • —Temporal action padding masked from loss. Original D1 normalization retained; new intervention statistics recorded separately.
  • —LeRobot 0.6.1, PyTorch 2.11.0+cu128, torchvision 0.26.0+cu128, Transformers 5.5.4, Accelerate 1.14.0. Eight RTX PRO 6000 Blackwell GPUs.

Inference export

Weights and normalization tensors are unchanged. The training-only mask_action_padding_loss config field is removed from the inference config for stock LeRobot compatibility; the original config is in reproducibility/. This does not alter inference computation. The tokenizer is bundled at repository root and its processor reference points to this repository instead of ephemeral training storage. For fully offline use, override tokenizer_processor.tokenizer_name to your downloaded local repository directory.

Policy expects three canonical RGB keys: observation.images.base_0_rgb, observation.images.left_wrist_0_rgb, and observation.images.right_wrist_0_rgb; 14-dimensional state/actions, 50-action prediction horizon. Apply the checkpoint processors and the appropriate observation rename map.

Verification and reproducibility

Strict weight reload, all-model finite checks, real-observation inference (1, 50, 14), and persistent checkpoint copy checksum passed. These are software integrity checks, not a robot success-rate evaluation.

Visible reproducibility/ files contain the original D1 selection and rank, intervention source manifest and segment mapping, statistics, frozen-code hashes and training patch, training command, scripts, logs and validation reports. Historical scripts contain machine-specific paths and require adaptation elsewhere. Optimizer state is deliberately not uploaded; this is an inference export, not an exact optimizer-resume bundle.

Model SHA256: dccc8e857ad9d3ae3ae4aac42925404a53aa667d98f1945461cd68899d620f9c.