datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pusht_96_norm2_remap
pusht_96_norm2
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2_remap.pusht_96_norm4
pusht_96_norm4
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
200
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4.pusht_96_norm4_10hz
pusht_96_norm4_10hz
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 10Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Each episode is success-only and capped at 150 environment steps.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_10hz.pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.pusht_96_norm4_visual_nomarker
pusht_96_norm4_visual_nomarker
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 90.
Each episode is success-only and capped at 30 environment steps.
Visual marker mode: none. The pusher is rendered with radius 11.0 in 512-space; physics still uses the environment's collision radius.
Prompt mode:… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker.pusht_96_norm4_remap
pusht_96_norm4
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
200
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_remap.pusht_96_norm4_10hz_remap
pusht_96_norm4_10hz
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 10Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Each episode is success-only and capped at 150 environment steps.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_10hz_remap.pusht_96_norm2
pusht_96_norm2
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2.repro-siamesenorm-breaking-the-barrier-to-reconciling-pre-post-norm-traces
Agent traces
Agent sessions published from a Trackio Logbook.
pusht_96_norm4_cot_chunk_k3_20260622_perseg
pusht_96_norm4_cot_chunk_k3_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k3_20260622_perseg.pusht_96_norm4_cot_chunk_k10_20260622_perseg
pusht_96_norm4_cot_chunk_k10_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k10_20260622_perseg.pusht_96_norm4_noncot_chunk_k3_20260622_perseg
pusht_96_norm4_noncot_chunk_k3_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_noncot_chunk_k3_20260622_perseg.pusht_96_norm4_cot_chunk_k1_20260622_perseg
pusht_96_norm4_cot_chunk_k1_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k1_20260622_perseg.pusht_96_norm4_noncot_chunk_k10_20260622_perseg
pusht_96_norm4_noncot_chunk_k10_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_noncot_chunk_k10_20260622_perseg.pusht_96_norm4_cot_chunk_kinf_branchcmp_20260709_perseg
pusht_96_norm4_cot_chunk_kinf_branchcmp_20260709_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout + NO-LEAK branching comparison (89.7% rows branched: at the max-coverage-jump key step, 4 shuffled candidates — expert + 3 strictly-worse Gaussian alts (σ=48px) — each imagined one step; winner named only in the conclusion)) for the
BAGEL-7B-MoT feedback-interval study.… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_kinf_branchcmp_20260709_perseg.pusht_96_norm4_noncot_chunk_k5_20260622_perseg
pusht_96_norm4_noncot_chunk_k5_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_noncot_chunk_k5_20260622_perseg.pikabu_text_normTexts inverse normalized obtained from pikabu dataset.
Normalized using these notebooks for a personal russian normalization model (avaliable on HF Space as well).
All put into single jsonl file with lines like (beautified):
{
"tn": "\\- Ну как так то? У нас в Норильске при минус сорока градусах в буран люди не замерзают, а у вас при минус десяти без ветра человек насмерть замёрз?",
"itn": "\\- Ну как так то? У нас в Норильске при минус 40 градусах в буран люди не замерзают, а у вас при… See the full description on the dataset page: https://huggingface.co/datasets/saarus72/pikabu_text_norm.pusht_96_norm4_cot_chunk_k5_20260622_perseg
pusht_96_norm4_cot_chunk_k5_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k5_20260622_perseg.pusht_96_norm4_cot_chunk_kinf_20260707_perseg
pusht_96_norm4_cot_chunk_kinf_20260707_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_kinf_20260707_perseg.normalize_symlinkspusht_norm4_stopreq_plain_aligned100k
PushT Plain Stopreq Aligned To CoT 100k
Plain records are selected from /data/home/jiaxin/unified_world_model/data/pusht_96_norm4_visual_nomarker_data/data by the CoT source_record manifest.
Images and actions are unchanged; only the full prompt receives the stop-required line.
ehristoforu__fq2.5-7b-it-normalize_false-details
Dataset Card for Evaluation run of ehristoforu/fq2.5-7b-it-normalize_false
Dataset automatically created during the evaluation run of model ehristoforu/fq2.5-7b-it-normalize_false
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__fq2.5-7b-it-normalize_false-details.ehristoforu__fq2.5-7b-it-normalize_true-details
Dataset Card for Evaluation run of ehristoforu/fq2.5-7b-it-normalize_true
Dataset automatically created during the evaluation run of model ehristoforu/fq2.5-7b-it-normalize_true
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__fq2.5-7b-it-normalize_true-details.pusht_norm4_stopreq_cot_aligned100k
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
legacy_pusht_norm4_allstep_cot_stopreq_aligned100k_ordered
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
pusht_norm4_stopreq_cot_singlenext_keystep_ablation
pusht_norm4_stopreq_cot_singlenext_keystep_ablation
BAGEL VLM-Gym SFT dataset (pusht cot ABLATION).
ABLATION: PushT all-step CoT, single-next key-step variant (matches the 14608 ablation SFT); ordered
local source dir: pusht_96_norm4_visual_nomarker_allstep_thinking_cot_stopreq_singlenext_keystep_ablation_20260607_ordered
format: trajectories/<task>/{train,test}/*.jsonl.gz (ordered by __bagel_order_key); base64 images inline.
plain & cot variants are aligned 1:1 by… See the full description on the dataset page: https://huggingface.co/datasets/novastar114/pusht_norm4_stopreq_cot_singlenext_keystep_ablation.marcuscedricridia__cursa-o1-7b-v1.2-normalize-false-details
Dataset Card for Evaluation run of marcuscedricridia/cursa-o1-7b-v1.2-normalize-false
Dataset automatically created during the evaluation run of model marcuscedricridia/cursa-o1-7b-v1.2-normalize-false
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/marcuscedricridia__cursa-o1-7b-v1.2-normalize-false-details.normalize_fixed8_rewrite
