ultrastar111/pusht_96_norm4_noncot_chunk_k3_20260622_perseg
pusht_96_norm4_noncot_chunk_k3_20260622_perseg PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_noncot_chunk_k3_20260622_perseg.
pusht96norm4noncotchunkk320260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study.
- Format: gzipped JSONL shards under
training/, 1 row = 1 packed episode. CoT rows: per-segment layout —<think>per-step imagined frame (MSE target)</think>+ committed action chunk, with a loss-0"Action executed." + real framere-grounding turn between chunks. nonCoT rows: bare action chunks. Build argv + git rev inmetadata/. - Companion checkpoint:
ultrastar112/pusht_96_norm4_noncot_chunk_k3_world_model_20260622_perseg - Full study docs:
0710_train.md/0711_mulnode_eval.md/0711_hf_release.mdin the study repo. - License: CC-BY-NC-4.0 (research use).
