always-on
agentboard-babyai-v1-v071-always-on-switch50
BabyAI v0.7.1 Always-on and Switch-50 Runs
This dataset archives ten Qwen/Qwen3.5-9B Prime-RL runs on the packaged
AgentBoard BabyAI v0.7.1 environment. It contains four always-on objectives and
six schedules that switch objectives at training step 50.
Runs
Always-on runs:
RL-only
ECHO 0.05
ECHO 0.5
ECHO 1.0
Step-50 switch runs:
RL 50, then ECHO 0.05
ECHO 0.05, then RL 50
RL 50, then ECHO 0.5
ECHO 0.5, then RL 50
RL 50, then ECHO 1.0
ECHO 1.0, then RL 50
All… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-always-on-switch50.9B_iter3_main_on_policy_but_delete_always_guessbabyai-qwen35-9b-v070-alwayson-200-run1
BabyAI Qwen3.5-9B: RL-only vs ECHO 1.0
This repository contains the complete artifacts from one pair of 200-step
BabyAI training runs comparing RL-only against always-on ECHO 1.0.
Experiment
Model: Qwen/Qwen3.5-9B
Trainer: Prime-RL v0.7.0 (d334ea52)
Verifiers: v0.2.0
Environment: agentboard-babyai-v1-context
Training split: 84 tasks (three examples per BabyAI subtask)
Held-out split: 28 tasks (one example per subtask)
Training steps: 200
Batch size: 128
Group… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/babyai-qwen35-9b-v070-alwayson-200-run1.
