CoolFace
Datasetpublic

ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg

sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes5downloads
Dataset Card

sokobaneasyv8cotchunkk10worldmodel20260622_perseg

Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study.

  • —Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See metadata/ for the exact build argv + git rev.
  • —Source: sokoban easy v8 (res 256) per-step episodes.
  • —Built with: scripts/data/remap_sokoban_*_chunk_* in the study repo; see 0710_train.md.
  • —Companion checkpoints: ultrastar112/... · License: CC-BY-NC-4.0 (research use).