datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
world_model_eval_logssokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg.sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg.sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg.sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg
sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg.sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg.world_model_evalsokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg.sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg.sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg.world_model_for_wa_qasokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg.
