datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libero-microwave-grm-rollouts
LIBERO Microwave GRM Rollouts
Dense-reward-annotated rollout dataset from GRPO training of OpenVLA-OFT on LIBERO-10 Task 9 ("put the yellow and white mug in the microwave and close it").
Dataset
Stat
Value
Episodes
~2,100 (T >= 5 steps)
Format
LeRobot (parquet + images)
Task
put the yellow and white mug in the microwave and close it
Size
41 GB
Reward model
Robo-Dopamine GRM-3B
Policy
OpenVLA-OFT (LoRA, GRPO-trained)
Simulator
LIBERO… See the full description on the dataset page: https://huggingface.co/datasets/Auryal/libero-microwave-grm-rollouts.B1k_Rollouts
B1k_Rollouts
BEHAVIOR-1K policy rollouts in a modality-first layout: each top-level folder is a data
modality, and every modality holds one task-NNNN/ subfolder per task. Adding a task means
dropping a task-NNNN/ directory into each of the six folders — no restructuring.
data/task-NNNN/ per-episode parquet
meta/task-NNNN/
episodes/ episode json + bddl_transitions json
predicate_catalogs/ per-instance tracked-predicate catalog
trajectories/task-NNNN/… See the full description on the dataset page: https://huggingface.co/datasets/fastwalker1118/B1k_Rollouts.ps4mas-final-test-rollouts-0813
PS4MAS Final Test Rollouts (0813)
Source split: ps4mas-0521-splits final_test_scenarios.jsonl
Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json.
Files
Model
Rows
Path
best_rl_gigpo_debate_step40
800
traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.feval-sn47-rolloutsninja-rollouts-polarDeprecated/Old data, new data lives at https://huggingface.co/datasets/Wejh/ninja-agent-traces
userlm_rl_experiment_rolloutslie-detection-rollouts
Lie Detection Rollouts
Assistant completions across many open-weight models on the lie-detection
evaluation suite used by the
deception research pipeline. One subset per model,
one split per task.
Columns
messages — list of OpenAI-style messages. Each message has:
role: system | user | assistant
content: final message text
reasoning_content: chain-of-thought for reasoning models, None otherwise
is_lie — ground-truth label from the is_deceptive scorer:
lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.roboprobe-gpt55-rollouts
RoboProbe GPT-5.5 Rollouts
Camera videos, raw planner traces, and generated viewer manifests used by
roboprobe-console.
Planner condition: gpt55
Cameras: head, left wrist, right wrist
Browser: companion Hugging Face Static Space
math-rollouts
Mathematical Reasoning Rollouts Dataset
This dataset contains step-by-step reasoning rollouts generated with DeepSeek R1-Distill language models solving mathematical problems from the MATH dataset.
The dataset is designed for analyzing reasoning patterns, branching factors, planning strategies, and the effectiveness of different reasoning approaches in mathematical problem-solving.
Dataset Details
Dataset Description
Curated by: Uzay Macar, Paul C.… See the full description on the dataset page: https://huggingface.co/datasets/uzaymacar/math-rollouts.isaac-gr00t-pld-eval-rolloutsmemory-rolloutsmisalignment-indicators-bloom-rolloutsMC-Math-Rollouts
MC-Math-Rollouts
Step-level Monte Carlo resampling rollouts of Qwen3 models on competition math benchmarks, for measuring the advantage and importance of individual reasoning steps.
🔎 Explore the data interactively: MC-Math-Rollouts Value Profiles viewer — browse per-step value profiles for individual traces without downloading anything.
For each seed response, the chain of thought is split into steps ("reasoning prefixes"). From the end of every prefix, the model is resampled… See the full description on the dataset page: https://huggingface.co/datasets/kducohere/MC-Math-Rollouts.watercolour-rollouts-judge-led
Watercolour rollouts, judge-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 861 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the run with the original reward mix from the write-up, where the pairwise judge and
its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.rollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.watercolour-rollouts-hps-only
Watercolour rollouts, HPS-only run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 470 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step.
The point of the dataset is that it holds the whole run, not the good bits. Step 0 and
step 59 are both here, with the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-only.gpt-oss-20b-rollouts
GPT-OSS-20B Rollouts
Generated rollouts from GPT-OSS-20B with parsed Harmony channels (assistant thinking/final).
Schema: user_content, system_reasoning_effort, assistant_thinking, assistant_content.
Loading example: load_dataset("andyrdt/gpt-oss-20b-rollouts", "HarmBench", split="standard_train").
Notes
This repository uses manual configuration to expose both subset (config) and split dropdowns in the viewer.
Safety and jailbreak
HarmBench: Safety prompts… See the full description on the dataset page: https://huggingface.co/datasets/andyrdt/gpt-oss-20b-rollouts.feval-sn47-rolloutsswift-reasoning-rollouts-deepscaler-ministral8b
DeepScaleR Reasoning Rollouts (Ministral-8B)
This dataset contains reasoning rollouts used to train the SWIFT reward head.
Paper page: https://huggingface.co/papers/2505.12225
GitHub: https://github.com/aster2024/SWIFT/
Generator model: mistralai/Ministral-8B-Instruct-2410 (https://huggingface.co/mistralai/Ministral-8B-Instruct-2410)
Dataset Description
This dataset contains 10000 samples corresponding to the Generalization Test setup.
Source: DeepScaleR.
Generator:… See the full description on the dataset page: https://huggingface.co/datasets/Aster2024/swift-reasoning-rollouts-deepscaler-ministral8b.screen-highlighter-4b-highlight-only-v1-rolloutsdipper_rolloutsmaxrl_qwen3_4B_base_polaris_rollouts
MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts)
Every training rollout from an online RL run, with exact token ids, sampling
log-probs, and raw rewards — usable as a replay buffer to study off-policy RL
for LLM reasoning completely offline.
The run: Qwen3-4B-Base trained with the maxRL advantage estimator
(A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt;
maxRL paper) and a pure REINFORCE loss
(L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.ViZDoom-Agentic-Rollouts
ViZDoom Agentic Rollouts
This dataset contains model-generated evaluation results for the pufanyi/ViZDoom benchmark. The benchmark repository contains only test specifications; this repository contains scores and rollout artifacts.
Both configs contain all 120 evaluated episodes: 10 seeds for each of the 12 default single-player ViZDoom environments. Browse the synchronized videos in the ViZDoom Agentic Demo.
Configs
qwen3.6-27b-5tics
The… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/ViZDoom-Agentic-Rollouts.moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.thinking-rollouts
thinking-rollouts
Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT
saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset
of temperature-sweep-data.
Hive-partitioned Parquet, thinking_mode folded into the model tag:
rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think,
qwen3-1.7b-nothink). 100 samples/instance.
Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.envace2.0-envscaler-grpo-32gpu-rollouts-20260520
EnvACE 2.0 EnvScaler GRPO 32-GPU Rollouts (2026-05-20)
This dataset is the run-artifact backup for envscaler_non_conv_rl_grpo_32gpu_aligned_20260520_055146, an aligned 32-GPU non-conversation GRPO training run.
Contents
rollouts/train/: 201 JSONL dumps, steps 0 through 200
rollouts/val/: 41 JSONL dumps, validation snapshots through step 200
logs/: 177 driver, actor, reference, reward, and environment logs
419 backed-up business files in total
The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/xuzishan/envace2.0-envscaler-grpo-32gpu-rollouts-20260520.screen-highlighter-2b-highlight-only-v1-rolloutsur5e_pi05_rollouts
UR5e pi0.5 rollouts
Policy rollouts of mahgoobi/ur5e_pi05_10k on a
real UR5e with a Robotiq gripper, task "put the cup in the bowl", collected 2026-09-01 from
the deployment page in RoboResearch
(roboresearch.evaluation.ur5e.ui). 25 episodes at 20 Hz, action chunks of 50 steps
with action_steps of them executed per chunk (25 for every episode but the first, which
ran 10). 20 successes, 5 failures. Every failure is in a scene with
three or more distractors; the clean and… See the full description on the dataset page: https://huggingface.co/datasets/mahgoobi/ur5e_pi05_rollouts.real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rollouts.rollout_smolvla_v5_lora64
