CoolFace
Modelpublic

RooibosT/gr00t-n1.7-g1-dex1-bct-relarm-aug-30hz-h40

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes17downloads
Model Card

GR00T N1.7 — Unitree G1 + Dex1, Building Children's Table (30 Hz, relative arms, augmented state)

Action-expert fine-tune of nvidia/GR00T-N1.7-3B for a Unitree G1 with Dex1 grippers assembling a children's table. Selected by open-loop scan over 6 checkpoints, not by eval_loss (see Checkpoint selection below).

Inputs

ModalityKeysNotes
video (3)cam_head, cam_left_wrist, cam_right_wrist4th head camera was ablated and dropped — it changed arm accuracy by 0.00%
state (46)legs(12) waist(3) left_arm(7) right_arm(7) left_gripper(1) right_gripper(1) base_gravity(3) left_eef(6) right_eef(6)see State conventions
languageannotation.human.task_description5 subtasks

Outputs

Chunk of 40 steps (1.33 s at 30 Hz) over 19 dims:

KeyDimsRepresentation
waist3ABSOLUTE
left_arm, right_arm7 + 7RELATIVE (delta from current state)
left_gripper, right_gripper1 + 1ABSOLUTE

Legs are observed but not commanded — locomotion stays with the whole-body controller.

State conventions (must be reproduced exactly at deployment)

  • —base_gravity — gravity direction in the pelvis frame, from the IMU quaternion (w,x,y,z): -[2(xz - wy), 2(yz + wx), 1 - 2(x² + y²)]. Yaw-invariant, unlike the raw quaternion whose w/z components are ~all session heading (yaw std 109° vs roll 0.6°).
  • —left_eef / right_eef — wrist pose recomputed by forward kinematics from the arm joints in the same row, not the recorded end-effector state (which lagged the joints by 10 frames / 333 ms). Convention: g1_body29_hand14.urdf, waist joints zeroed (torso frame), wrist_yaw origin translated +0.05 m along local x, orientation as extrinsic xyz Euler.
  • —Root x/y/z are deliberately absent: x,y vary ~90x more between clips than within one (a clip-ID shortcut) and z is a deterministic function of the leg joints already in state.

A mismatch in these conventions silently corrupts 15 of 46 input dims. state_dropout=0.2 makes the model robust to missing state, not to wrong state.

Open-loop accuracy (66 held-out episodes, 529 windows, unnormalized units)

Executed stepsRe-inference periodArm MAEGripper MAE
50.17 s1.20°0.080
80.27 s1.50°0.091
160.53 s2.25°0.122
40 (full chunk)1.33 s3.98°0.198

Error grows with horizon, so the full-chunk figure is not the deployment figure — execute the head of the chunk and re-infer. A hold-current-pose baseline gives 1.13° at h1 and 5.72° over the full chunk. For scale, the arms move 10.2° over a 40-step chunk.

Training

Data1,249 train clips / 500,038 frames @ 30 fps, teleop stalls removed, subtasks rebalanced
Trainableaction expert only (1.62B of 3.14B); Cosmos-Reason2 backbone frozen
Schedule15,000 steps, effective batch 192, cosine, lr 1e-4, warmup 0.05, wd 1e-5
Regularizationstate_dropout 0.2
Hardware3x A100 80GB, DDP with bf16 gradient communication, 12h22m

Checkpoint selection

eval_loss is the flow-matching regression objective at random noise levels; across five runs it rose after ~7.5k steps while open-loop action accuracy kept improving. It is not a model-selection signal here. Checkpoints were instead scanned with a fixed-seed open-loop evaluation on the held-out split.

A longer schedule was tried and rejected: 30,000 steps at effective batch 256 saturated by ~20k and finished slightly worse than this model (full-chunk MSE 0.0299 vs 0.0291, gripper 0.204 vs 0.198), so roughly 5-6 epochs is enough for this dataset.