CoolFace
Datasetpublic

ultrastar111/maze2d_easy_native256_noncot_chunk_k3_20260707_perseg

maze2d_easy_native256_noncot_chunk_k3_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k3_20260707_perseg.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes26downloads
Dataset Card

maze2deasynative256noncotchunkk320260707_perseg

Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study.

  • —Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real frame re-grounding turn between chunks. nonCoT rows: bare action chunks. Build argv + git rev in metadata/.
  • —Companion checkpoint: ultrastar112/maze2d_easy_native256_noncot_chunk_k3_world_model_20260707_perseg
  • —Full study docs: 0710_train.md / 0711_mulnode_eval.md / 0711_hf_release.md in the study repo.
  • —License: CC-BY-NC-4.0 (research use).