sweagent/iter2-rl-rollouts
iter-2 RL rollouts (combo_fb, Qwen3.5-35B-A3B) Complete rollout + reward record for the iter-2 GRPO run: every trajectory the policy generated during training, with its graded reward. Preserved so the run stays re-analysable after its torch_dist checkpoints were retired (only latest-3 survive; the 11 HF milestones at iter_0/4/9/.../44/49 are the durable checkpoint record). Run base model Qwen3.5-35B-A3B init iter_49 of the iter-1 RL run… See the full description on the dataset page: https://huggingface.co/datasets/sweagent/iter2-rl-rollouts.
036
iter-2 RL rollouts: 20,248 training (23,139 trajs) + 5,330 eval, 50 GRPO steps
initial commit
