CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Auryal /libero-microwave-grm-rollouts LIBERO Microwave GRM Rollouts Dense-reward-annotated rollout dataset from GRPO training of OpenVLA-OFT on LIBERO-10 Task 9 ("put the yellow and white mug in the microwave and close it"). Dataset Stat Value Episodes ~2,100 (T >= 5 steps) Format LeRobot (parquet + images) Task put the yellow and white mug in the microwave and close it Size 41 GB Reward model Robo-Dopamine GRM-3B Policy OpenVLA-OFT (LoRA, GRPO-trained) Simulator LIBERO… See the full description on the dataset page: https://huggingface.co/datasets/Auryal/libero-microwave-grm-rollouts.robotics1K<n<10K0 likes7.6k downloads7mo agoHugging Face02bag100 /action-atlas-rollout-videos Action Atlas — VLA Rollout Videos Local rollout/ablation videos for Pi0.5, OpenVLA-OFT, X-VLA, GR00T, SmolVLA, and ACT/ALOHA, organized by model. Companion to action-atlas-{pi05,oft,xvla,groot,smolvla} (SAEs + activations + concepts). 368283 unique mp4 clips, 61.8 GB. Per-model: {'act_aloha': 990, 'groot': 163891, 'oft': 24284, 'pi05': 62468, 'smolvla': 56844, 'xvla': 59806} manifest.jsonl: one row per clip (model, env, experiment, sha256, bytes, hf_path). 0 likes6k downloads3mo agoHugging Face03fastwalker1118 /B1k_Rollouts B1k_Rollouts BEHAVIOR-1K policy rollouts in a modality-first layout: each top-level folder is a data modality, and every modality holds one task-NNNN/ subfolder per task. Adding a task means dropping a task-NNNN/ directory into each of the six folders — no restructuring. data/task-NNNN/ per-episode parquet meta/task-NNNN/ episodes/ episode json + bddl_transitions json predicate_catalogs/ per-instance tracked-predicate catalog trajectories/task-NNNN/… See the full description on the dataset page: https://huggingface.co/datasets/fastwalker1118/B1k_Rollouts.robotics0 likes4.8k downloads23d agoHugging Face04yinita /ps4mas-final-test-rollouts-0813 PS4MAS Final Test Rollouts (0813) Source split: ps4mas-0521-splits final_test_scenarios.jsonl Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json. Files Model Rows Path best_rl_gigpo_debate_step40 800 traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.tabular1K<n<10K0 likes4.5k downloads7d agoHugging Face05Wejh /ninja-rollouts-polarDeprecated/Old data, new data lives at https://huggingface.co/datasets/Wejh/ninja-agent-traces 3 likes4.1k downloads2mo agoHugging Face06thermopylae-777 /feval-sn47-rollouts0 likes4k downloads15d agoHugging Face07fastwalker1118 /B1k_rollout B1k_rollout — Pi0.5 policy rollouts on BEHAVIOR-1K (2025 challenge) ⚠️ RULE: never roll out B1K official test instances (test-set contamination) Each task's metadata/test_instances.csv row = official public-test ids (dense idx 0–19) FOLLOWED BY our generated ids (the contiguous 501+/701+ block, dense idx ≥20). The first entries (dense idx 0–19) are B1K's official evaluation instances — rolling them out and training on them is test-set contamination. Only ever roll… See the full description on the dataset page: https://huggingface.co/datasets/fastwalker1118/B1k_rollout.1 likes3.5k downloads2mo agoHugging Face08momergul /userlm_rl_experiment_rolloutsgated0 likes3.3k downloads1mo agoHugging Face09ai-safety-institute /lie-detection-rollouts Lie Detection Rollouts Assistant completions across many open-weight models on the lie-detection evaluation suite used by the deception research pipeline. One subset per model, one split per task. Columns messages — list of OpenAI-style messages. Each message has: role: system | user | assistant content: final message text reasoning_content: chain-of-thought for reasoning models, None otherwise is_lie — ground-truth label from the is_deceptive scorer: lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.text1M<n<10M0 likes2.8k downloads3mo agoHugging Face10Solomonz /roboprobe-gpt55-rollouts RoboProbe GPT-5.5 Rollouts Camera videos, raw planner traces, and generated viewer manifests used by roboprobe-console. Planner condition: gpt55 Cameras: head, left wrist, right wrist Browser: companion Hugging Face Static Space videorobotics1K<n<10K0 likes2.4k downloads10d agoHugging Face11uzaymacar /math-rollouts Mathematical Reasoning Rollouts Dataset This dataset contains step-by-step reasoning rollouts generated with DeepSeek R1-Distill language models solving mathematical problems from the MATH dataset. The dataset is designed for analyzing reasoning patterns, branching factors, planning strategies, and the effectiveness of different reasoning approaches in mathematical problem-solving. Dataset Details Dataset Description Curated by: Uzay Macar, Paul C.… See the full description on the dataset page: https://huggingface.co/datasets/uzaymacar/math-rollouts.textquestion-answering10K<n<100K11 likes2.2k downloads1y agoHugging Face12insagur /isaac-gr00t-pld-eval-rollouts0 likes1.6k downloads6mo agoHugging Face13kzhou35 /misalignment-indicators-bloom-rollouts0 likes1.5k downloads3mo agoHugging Face14latency-sensitive-bench /memory-rolloutsimage10M<n<100M0 likes1.5k downloads7d agoHugging Face15FineEnvs /watercolour-rollouts-judge-led Watercolour rollouts, judge-led run Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one. Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by writing p5.brush sketches. 861 paintings, the sketch that produced each one, and the reward it earned, indexed by training step. This is the run with the original reward mix from the write-up, where the pairwise judge and its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.imagetext-to-imagen<1K0 likes1.4k downloads21d agoHugging Face16ryanjin333 /behavior-pi05-all100-rollouts BEHAVIOR pi0.5 all-100 rollouts Public storage for videos and per-episode metrics from the all-100 experiment. Evaluation and replay run on the existing Linux machine, not Nebius. No trained all-100 policy evaluation has been published yet. Every result must identify its policy commit, task, instance, seed, success, Q-score, steps and termination reason. Training replay and evaluation episodes remain separate. Experiment plan and policies: behavior-pi05-all100. videon<1K0 likes1.4k downloads12m agoHugging Face17kducohere /MC-Math-Rollouts MC-Math-Rollouts Step-level Monte Carlo resampling rollouts of Qwen3 models on competition math benchmarks, for measuring the advantage and importance of individual reasoning steps. 🔎 Explore the data interactively: MC-Math-Rollouts Value Profiles viewer — browse per-step value profiles for individual traces without downloading anything. For each seed response, the chain of thought is split into steps ("reasoning prefixes"). From the end of every prefix, the model is resampled… See the full description on the dataset page: https://huggingface.co/datasets/kducohere/MC-Math-Rollouts.text-generation0 likes1.3k downloads2mo agoHugging Face18Heisen0928 /robotwin_rollouttabular100K<n<1M0 likes1k downloads7mo agoHugging Face19kylemontgomery /imo-1030-rollout-partial0 likes1k downloads11mo agoHugging Face20FineEnvs /watercolour-rollouts-hps-only Watercolour rollouts, HPS-only run Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one. Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by writing p5.brush sketches. 470 paintings, the sketch that produced each one, and the reward it earned, indexed by training step. The point of the dataset is that it holds the whole run, not the good bits. Step 0 and step 59 are both here, with the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-only.imagetext-to-imagen<1K0 likes973 downloads21d agoHugging Face21nasa-cisto-data-science-group /aurora_rollout_beta0 likes959 downloads2y agoHugging Face22reasoning-cues /rollouts-olmo7b-cue-search rollouts-olmo7b-cue-search Model: allenai/Olmo-3-1025-7B (snapshot a81bae42). Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42). Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms. Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.tabular100K<n<1M0 likes883 downloads10d agoHugging Face23andyrdt /gpt-oss-20b-rollouts GPT-OSS-20B Rollouts Generated rollouts from GPT-OSS-20B with parsed Harmony channels (assistant thinking/final). Schema: user_content, system_reasoning_effort, assistant_thinking, assistant_content. Loading example: load_dataset("andyrdt/gpt-oss-20b-rollouts", "HarmBench", split="standard_train"). Notes This repository uses manual configuration to expose both subset (config) and split dropdowns in the viewer. Safety and jailbreak HarmBench: Safety prompts… See the full description on the dataset page: https://huggingface.co/datasets/andyrdt/gpt-oss-20b-rollouts.text1M<n<10M7 likes709 downloads9mo agoHugging Face24YuMoool /astra-robodojo-rollouts Astra RoboDojo Evaluation Records Rollout records from the evaluations in GPT 6 Astra as an Embodied Policy, by Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. This archive includes action proposals, executed actions, observations, robot states, model-provided explanations and reasoning summaries, and metadata for reproducing the evaluation settings, together with a reader and documentation. Report Public controller source Data schema and alignment… See the full description on the dataset page: https://huggingface.co/datasets/YuMoool/astra-robodojo-rollouts.imageroboticsn<1K0 likes705 downloads8d agoHugging Face25Christelle04 /rollout_act_v5tabular100K<n<1M0 likes689 downloads1mo agoHugging Face26GohaDAYN /feval-sn47-rollouts0 likes665 downloads9d agoHugging Face27mustafaah /screen-highlighter-4b-highlight-only-v1-rolloutsimage10K<n<100K0 likes662 downloads10d agoHugging Face28pufanyi /ViZDoom-Agentic-Rollouts ViZDoom Agentic Rollouts This dataset contains model-generated evaluation results for the pufanyi/ViZDoom benchmark. The benchmark repository contains only test specifications; this repository contains scores and rollout artifacts. Both configs contain all 120 evaluated episodes: 10 seeds for each of the 12 default single-player ViZDoom environments. Browse the synchronized videos in the ViZDoom Agentic Demo. Configs qwen3.6-27b-5tics The… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/ViZDoom-Agentic-Rollouts.tabularreinforcement-learningn<1K0 likes625 downloads2mo agoHugging Face29samuki-hf /thinking-rollouts thinking-rollouts Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset of temperature-sweep-data. Hive-partitioned Parquet, thinking_mode folded into the model tag: rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think, qwen3-1.7b-nothink). 100 samples/instance. Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.tabulartext-generation10M<n<100M2 likes619 downloads2mo agoHugging Face30ftajwar /maxrl_qwen3_4B_base_polaris_rollouts MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts) Every training rollout from an online RL run, with exact token ids, sampling log-probs, and raw rewards — usable as a replay buffer to study off-policy RL for LLM reasoning completely offline. The run: Qwen3-4B-Base trained with the maxRL advantage estimator (A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt; maxRL paper) and a pure REINFORCE loss (L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.tabulartext-generation1M<n<10M0 likes598 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.