CoolFace
Apppublic

WitneyWW/acts-causal-rollouts

sourceHugging Faceupdated 8h agoView on Hugging Face
0likes
App README

ACTS causal rollouts

Motherboard (epochs 001, 004, 005, 010), PushT (epochs 003, 007, 013), and Rope (epochs 001, 005) held-out contact-stratified predictions, plus Motherboard custom-attention epochs 002, 004 and 010: 268 short and 264 long videos. All Motherboard checkpoints share the same 20 selected windows; both Rope checkpoints share 34. Sampling settings and seeds match. Motherboard · custom attn (nanoactioncausalcustomv1, epoch 002) uses the same data, split, recipe, windows and seeds as the other Motherboard checkpoints (nanoactioncausalumiallv1); only stream routing differs. The other runs let every stream attend every stream (time-causal only). The custom run restricts attention per modality: video reads wrists, tactile and view actions; each wrist reads video, its own-hand tactile and view action; each tactile stream reads video, its own-hand wrist and sensor action; left/right wrists and left/right tactile never attend each other. Its epoch (002) is not matched to 001/005/010. In 9 of its 20 long rollouts all five streams diverge to noise between generated windows 11 and 14; all nine are in the 0909 session, and none of the eight 0911/0912 cases diverge. The all-stream checkpoints diverge in at most 1 of 20. Epoch 004 is the matched pair: both runs are at epoch 004 / step 44615 on the same data, split, recipe, windows and seeds, so they isolate stream routing alone. Neither arm diverges in any of the 20 long rollouts at this epoch, so the custom epoch 002 divergences reflect that earlier checkpoint rather than the routing. Averaged over the 20 paired cases, custom routing is slightly behind: main view -0.34 dB PSNR short and -0.51 dB long, right wrist -0.56/-0.92 dB, right tactile -1.33 dB long (paired t-test p<0.05 for those three long-horizon streams); left tactile is within noise. Custom epoch 010 (step 98153) is matched to all-streams epoch 010. Again no long rollout diverges in either arm, and custom stays slightly behind: main view -0.22 dB short / -0.09 dB long (not significant), right wrist -0.89/-0.72 dB and right tactile -1.29 dB long (p<0.05). Epoch 010 repeats the matched comparison with custom epoch 010 against all-streams epoch 010 (both at step 98153). Neither arm diverges in any of the 20 long rollouts. The main-view gap shrinks: -0.22 dB PSNR short and -0.09 dB long, neither significant. The custom run is still behind on the right-hand streams, with right wrist -0.89 dB short and -0.72 dB long, and right tactile -1.29 dB long (paired t-test p<0.05). PushT epochs 003 and 007 share 15 short and 14 long selections on four modelv3 test segments. Epoch 013 preserves those selections and adds the remaining four segments from its own mixedpushtwristholdout.json test split, giving 30 short and 28 long cases. Aggregate PushT scores across these different selections are not directly comparable. Rope uses nine test segments from mixedropewristholdout.json, with 34 short and 34 long cases. Two segments lack no-contact windows. Short: eight future frames. Long: 120 future frames at 15 FPS, using generated observation history and recorded actions. Five observation streams, GT/prediction videos, and PSNR/SSIM/LPIPS metrics.

Contact classes describe the initial window. Selection is qualitative and stratified, not a full-test-set benchmark. GT images are decoded cached latents. Metrics score future frames only; long metrics average over 15 generated windows. The tasks use different datasets/checkpoints and their scores are not a controlled model comparison. The earlier PushT modelv3 split resolves old episode006 names to current episode007. PushT 005seg01 has no full-contact case and a short-only onset. In epoch 013, 010seg02 also lacks full contact and 010seg03 has a short-only full-contact case. Motherboard 0909 force calibration differs from 0911/0912; a common 0.5 N contact threshold is used for selection.

This repository contains inference results only, not model checkpoints or training data. Web videos may be compressed for viewing; metrics use the original decoded predictions before video compression. Each sample includes recorded left/right normal force in newtons, synchronized to video frames. These are supplied recording measurements, not predicted forces. The 0.5 N contact threshold and history/future boundary are marked.