CoolFace
Datasetpublic

sweagent/iter2-rl-rollouts

iter-2 RL rollouts (combo_fb, Qwen3.5-35B-A3B) Complete rollout + reward record for the iter-2 GRPO run: every trajectory the policy generated during training, with its graded reward. Preserved so the run stays re-analysable after its torch_dist checkpoints were retired (only latest-3 survive; the 11 HF milestones at iter_0/4/9/.../44/49 are the durable checkpoint record). Run base model Qwen3.5-35B-A3B init iter_49 of the iter-1 RL run… See the full description on the dataset page: https://huggingface.co/datasets/sweagent/iter2-rl-rollouts.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes36downloads
settings

This repository belongs to sweagent on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameiter2-rl-rollouts
visibilitypublic
licencemit
gatedno
ownersweagent
Account settings
sweagent/iter2-rl-rollouts · CoolFace