ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg.
sokobaneasyv8cotchunkk10worldmodel20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study.
- Format: gzipped JSONL shards under
training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout:<think>per-step imagined frame (MSE target)</think>then the committed action chunk; between chunks a loss-0"Action executed." + real frame(GT re-grounding). Seemetadata/for the exact build argv + git rev. - Source: sokoban easy v8 (res 256) per-step episodes.
- Built with:
scripts/data/remap_sokoban_*_chunk_*in the study repo; see0710_train.md. - Companion checkpoints:
ultrastar112/...· License: CC-BY-NC-4.0 (research use).
