rollouts
Datasets
All datasets matching “rollouts”libero-microwave-grm-rollouts
LIBERO Microwave GRM Rollouts
Dense-reward-annotated rollout dataset from GRPO training of OpenVLA-OFT on LIBERO-10 Task 9 ("put the yellow and white mug in the microwave and close it").
Dataset
Stat
Value
Episodes
~2,100 (T >= 5 steps)
Format
LeRobot (parquet + images)
Task
put the yellow and white mug in the microwave and close it
Size
41 GB
Reward model
Robo-Dopamine GRM-3B
Policy
OpenVLA-OFT (LoRA, GRPO-trained)
Simulator
LIBERO… See the full description on the dataset page: https://huggingface.co/datasets/Auryal/libero-microwave-grm-rollouts.B1k_Rollouts
B1k_Rollouts
BEHAVIOR-1K policy rollouts in a modality-first layout: each top-level folder is a data
modality, and every modality holds one task-NNNN/ subfolder per task. Adding a task means
dropping a task-NNNN/ directory into each of the six folders — no restructuring.
data/task-NNNN/ per-episode parquet
meta/task-NNNN/
episodes/ episode json + bddl_transitions json
predicate_catalogs/ per-instance tracked-predicate catalog
trajectories/task-NNNN/… See the full description on the dataset page: https://huggingface.co/datasets/fastwalker1118/B1k_Rollouts.ps4mas-final-test-rollouts-0813
PS4MAS Final Test Rollouts (0813)
Source split: ps4mas-0521-splits final_test_scenarios.jsonl
Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json.
Files
Model
Rows
Path
best_rl_gigpo_debate_step40
800
traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.feval-sn47-rolloutsninja-rollouts-polarDeprecated/Old data, new data lives at https://huggingface.co/datasets/Wejh/ninja-agent-traces
userlm_rl_experiment_rollouts
