sweagent/iter2-rl-rollouts
iter-2 RL rollouts (combo_fb, Qwen3.5-35B-A3B) Complete rollout + reward record for the iter-2 GRPO run: every trajectory the policy generated during training, with its graded reward. Preserved so the run stays re-analysable after its torch_dist checkpoints were retired (only latest-3 survive; the 11 HF milestones at iter_0/4/9/.../44/49 are the durable checkpoint record). Run base model Qwen3.5-35B-A3B init iter_49 of the iter-1 RL run… See the full description on the dataset page: https://huggingface.co/datasets/sweagent/iter2-rl-rollouts.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face