CoolFace
Datasetpublic

bhoy/agentboard-babyai-v1-v071-always-on-switch50

BabyAI v0.7.1 Always-on and Switch-50 Runs This dataset archives ten Qwen/Qwen3.5-9B Prime-RL runs on the packaged AgentBoard BabyAI v0.7.1 environment. It contains four always-on objectives and six schedules that switch objectives at training step 50. Runs Always-on runs: RL-only ECHO 0.05 ECHO 0.5 ECHO 1.0 Step-50 switch runs: RL 50, then ECHO 0.05 ECHO 0.05, then RL 50 RL 50, then ECHO 0.5 ECHO 0.5, then RL 50 RL 50, then ECHO 1.0 ECHO 1.0, then RL 50 All… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-always-on-switch50.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes39downloads
Dataset Card

BabyAI v0.7.1 Always-on and Switch-50 Runs

This dataset archives ten Qwen/Qwen3.5-9B Prime-RL runs on the packaged AgentBoard BabyAI v0.7.1 environment. It contains four always-on objectives and six schedules that switch objectives at training step 50.

Runs

Always-on runs:

  • —RL-only
  • —ECHO 0.05
  • —ECHO 0.5
  • —ECHO 1.0

Step-50 switch runs:

  • —RL 50, then ECHO 0.05
  • —ECHO 0.05, then RL 50
  • —RL 50, then ECHO 0.5
  • —ECHO 0.5, then RL 50
  • —RL 50, then ECHO 1.0
  • —ECHO 1.0, then RL 50

All runs used 100 training steps, Qwen/Qwen3.5-9B, a training batch size of 128, eight rollouts per task group, a 20-turn environment cap, temperature 0.7, and thinking disabled. The held-out evaluation split contains 28 tasks. Each checkpoint evaluation uses three rollouts per task.

Layout

text
runs/
  always_on/<run>/
    adapters/step_<N>/
    logs/
    configs/
  switch50/<run>/
    adapters/step_<N>/
    logs/
    configs/
evals/
  always_on_rl_only_and_echo_0.05/
  always_on_echo_0.5_and_echo_1.0/
  switch50_all_weights/
training_samples/
analysis/
metadata/

Each adapters directory contains every saved LoRA broadcast: steps 5 through 95 at five-step intervals, plus steps 99 and 100. Each adapter includes adapter_model.safetensors, adapter_config.json, and the Prime-RL STABLE marker.

logs contains the Prime-RL inference, orchestrator, and trainer logs, plus a curated training_history.log with the complete steps 1-100 console history. The history log preserves per-step reward, candidate/trainable rollout counts, mean turns, truncation, off-policy, errors, and elapsed time. Launcher scripts are intentionally excluded. configs contains the resolved inference, orchestrator, and trainer TOMLs that Prime-RL actually used. TOML is a human-readable configuration format; these files record model, batching, sampling, optimizer, environment, and ECHO settings for reproducibility.

The evals folders contain per-checkpoint evaluation TOMLs, driver logs, and human-readable traces.jsonl files. analysis contains figures and aggregate CSVs. analysis/training_behavior_metrics.csv contains 100-step behavior curves for all ten runs, separately aggregated over all generated candidates and retained trainable rollouts. Its fields include reward distributions, success, turns, ran-out-of-turns rate, grounding, invalid actions, action repetition, token usage, generation time, policy versions, and errors. metadata/files.tsv records every archived file and its byte size.

training_samples contains 120 complete training trajectories selected from steps 10, 35, 65, and 95 across all ten runs. Each run/step includes a successful or high-reward trajectory, a median-reward trajectory, and a failed trajectory that ran out of turns. The combined raw JSONL preserves messages, token IDs, masks, log probabilities, rewards, metrics, timing, and policy metadata. Readable Markdown renderings and a CSV manifest are also included. This is a deterministic qualitative sample, not a random statistical sample.

Scope

Raw training rollout tensors and full optimizer checkpoints are intentionally excluded. The LoRA adapters are sufficient for inference and post-hoc evaluation, but they do not provide exact optimizer-state training resumption.

Always-on and switch schedules are independent training runs executed at different times and, in some cases, on different GPU nodes. Differences before the switch should not be interpreted as controlled switch effects.