CoolFace
Datasetpublic

justintiensmith/Spa_Bench_Partial_GR00T-N1.7_Full_Fine-Tune

Spa-Bench physical rollouts — GR00T-N1.7 Full Fine-Tune (partial) Physical SO-101 rollout recordings for the epoch-12 GR00T-N1.7 Full Fine-Tune Spa-Bench checkpoint. Scope and counts This is a partial familiar/in-distribution baseline, not a complete Spa-Bench evaluation. It contains 120 episodes—20 from each of the six task families—plus 127,143 frames, 90 unique instruction strings, two 480×640 RGB streams, and six-dimensional state/action at 30 FPS. It contains… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Spa_Bench_Partial_GR00T-N1.7_Full_Fine-Tune.

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
0likes334downloads
Dataset Card

Spa-Bench physical rollouts — GR00T-N1.7 Full Fine-Tune (partial)

Physical SO-101 rollout recordings for the epoch-12 GR00T-N1.7 Full Fine-Tune Spa-Bench checkpoint.

Scope and counts

This is a partial familiar/in-distribution baseline, not a complete Spa-Bench evaluation. It contains 120 episodes—20 from each of the six task families—plus 127,143 frames, 90 unique instruction strings, two 480×640 RGB streams, and six-dimensional state/action at 30 FPS. It contains no OOD or additional diagnostic trials.

The checkpoint completed 25/120 trials (20.8%). Because coverage is partial, it is excluded from the report's full OOD and diagnostic comparisons.

Non-gripper deployment actions used the causal filter filtered = 0.25 × current + 0.75 × previous_filtered, initialized from the measured state and reset for each rollout. This condition also received a small upward initialization assist before the scored timer; the linked model card records the retrospective displacement estimates.

Outcome labels and groupings are maintained in the Spa-Bench evaluation spreadsheets, not as a frame-level success field in this dataset.

Structure and use

Each trajectory stores the instruction/task index, six absolute SO-101 joint positions for state and action, and synchronized middle/wrist video. This release supports audit and qualitative failure analysis. It must not be used as evidence of OOD performance, and the initialization intervention prevents a clean comparison with the Frozen LLM configuration.

Do not train on these trajectories and report them as held-out performance. Robot deployment requires safeguards and human supervision.

Citation

Please cite the completed Spa-Bench MSc report and the thesis artifact.