datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_kinf_20260707_perseg.maze2d_easy_native256_noncot_chunk_k3_20260707_perseg
maze2d_easy_native256_noncot_chunk_k3_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k3_20260707_perseg.maze2d_easy_native256_cot_chunk_k5_20260707_perseg
maze2d_easy_native256_cot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k5_20260707_perseg.maze2d_easy_cot_chunk_k3_train
maze2d_easy_cot_chunk_k3_train
BAGEL VLM-Gym world-model dataset (maze2d / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=3 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_cot_chunk_k3_train.maze2d_easy_cot_chunk_k5_train
maze2d_easy_cot_chunk_k5_train
BAGEL VLM-Gym world-model dataset (maze2d / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=5 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_cot_chunk_k5_train.maze2d_easy_noncot_chunk_k10_train
maze2d_easy_noncot_chunk_k10_train
BAGEL VLM-Gym world-model dataset (maze2d / noncot).
Non-CoT chunk-K train set (no imagined reasoning); re-grounds every K=10 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT
variants share the same… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_noncot_chunk_k10_train.maze2d_easy_native256_cot_chunk_k10_20260707_perseg
maze2d_easy_native256_cot_chunk_k10_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k10_20260707_perseg.maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k10_20260707_perseg.maze2d_easy_cot_chunk_k10_train
maze2d_easy_cot_chunk_k10_train
BAGEL VLM-Gym world-model dataset (maze2d / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_cot_chunk_k10_train.maze2d_easy_noncot_chunk_kinf_train
maze2d_easy_noncot_chunk_kinf_train
BAGEL VLM-Gym world-model dataset (maze2d / noncot).
Non-CoT chunk-K train set (no imagined reasoning); re-grounds every K=inf (open-loop; imagine the whole episode) steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org;… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_noncot_chunk_kinf_train.maze2d_easy_cot_chunk_kinf_train
maze2d_easy_cot_chunk_kinf_train
BAGEL VLM-Gym world-model dataset (maze2d / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=inf (open-loop; imagine the whole episode) steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_cot_chunk_kinf_train.maze2d_easy_native256_noncot_chunk_k1_20260707_perseg
maze2d_easy_native256_noncot_chunk_k1_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k1_20260707_perseg.maze2d_easy_native256_noncot_chunk_k5_20260707_perseg
maze2d_easy_native256_noncot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k5_20260707_perseg.maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg
maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg.maze2d_easy_cot_chunk_k1_train
maze2d_easy_cot_chunk_k1_train
BAGEL VLM-Gym world-model dataset (maze2d / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=1 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_cot_chunk_k1_train.maze2d_easy_native256_cot_chunk_k1_20260707_perseg
maze2d_easy_native256_cot_chunk_k1_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k1_20260707_perseg.maze2d_easy_noncot_chunk_k1_train
maze2d_easy_noncot_chunk_k1_train
BAGEL VLM-Gym world-model dataset (maze2d / noncot).
Non-CoT chunk-K train set (no imagined reasoning); re-grounds every K=1 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT
variants share the same… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_noncot_chunk_k1_train.maze2d_easy_noncot_chunk_k3_train
maze2d_easy_noncot_chunk_k3_train
BAGEL VLM-Gym world-model dataset (maze2d / noncot).
Non-CoT chunk-K train set (no imagined reasoning); re-grounds every K=3 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT
variants share the same… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_noncot_chunk_k3_train.
