CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ryukijano /repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes104 downloads2mo agoHugging Face02drvp /world_model_eval_logstabularn<1K0 likes30 downloads2mo agoHugging Face03ultrastar111 /sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes17 downloads2mo agoHugging Face04ultrastar111 /sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes16 downloads2mo agoHugging Face05ultrastar111 /sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes14 downloads2mo agoHugging Face06ultrastar111 /sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg.tabularreinforcement-learning100K<n<1M0 likes13 downloads2mo agoHugging Face07ultrastar111 /sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes9 downloads2mo agoHugging Face08drvp /world_model_evaltabularn<1K0 likes8 downloads4mo agoHugging Face09ultrastar111 /sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes7 downloads2mo agoHugging Face10ultrastar111 /sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes7 downloads2mo agoHugging Face11ultrastar111 /sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes6 downloads2mo agoHugging Face12LangAGI-Lab /world_model_for_wa_qatabular10K<n<100K0 likes5 downloads2y agoHugging Face13ultrastar111 /sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.