willamazon1/Qwen3.5-9B-smith-v5-gdpo-exact
Qwen3.5-9B Smith AENV v5 (GDPO, exact reward) — RL checkpoint series
Reinforcement-learning checkpoint series from the `cpo_smith ... smith-v5-gdpo-exact` run: asynchronous multi-turn agentic-environment RL on Qwen/Qwen3.5-9B with an exact (verifiable) reward.
The policy was initialized from an internal SFT of Qwen/Qwen3.5-9B (qwen35_9b_sft_v3), which also served as the reference model.
Checkpoints
92 checkpoints, saved every 4 iterations up to 31, then every 2 iterations, from iter_0000003 to iter_0000199. Each lives in its own subfolder of this repo so you can compare points along the training curve:
iter_0000003, iter_0000007, iter_0000011, iter_0000015, iter_0000019, iter_0000023, iter_0000027, iter_0000031, iter_0000033, iter_0000035, iter_0000037, iter_0000039, iter_0000041, iter_0000043, iter_0000045, iter_0000047, iter_0000049, iter_0000051, iter_0000053, iter_0000055, iter_0000057, iter_0000059, iter_0000061, iter_0000063, iter_0000065, iter_0000067, iter_0000069, iter_0000071, iter_0000073, iter_0000075, iter_0000077, iter_0000079, iter_0000081, iter_0000083, iter_0000085, iter_0000087, iter_0000089, iter_0000091, iter_0000093, iter_0000095, iter_0000097, iter_0000099, iter_0000101, iter_0000103, iter_0000105, iter_0000107, iter_0000109, iter_0000111, iter_0000113, iter_0000115, iter_0000117, iter_0000119, iter_0000121, iter_0000123, iter_0000125, iter_0000127, iter_0000129, iter_0000131, iter_0000133, iter_0000135, iter_0000137, iter_0000139, iter_0000141, iter_0000143, iter_0000145, iter_0000147, iter_0000149, iter_0000151, iter_0000153, iter_0000155, iter_0000157, iter_0000159, iter_0000161, iter_0000163, iter_0000165, iter_0000167, iter_0000169, iter_0000171, iter_0000173, iter_0000175, iter_0000177, iter_0000179, iter_0000181, iter_0000183, iter_0000185, iter_0000187, iter_0000189, iter_0000191, iter_0000193, iter_0000195, iter_0000197, iter_0000199
Usage
Each training iteration is a subfolder of this repo, so pass subfolder= when loading:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "willamazon1/Qwen3.5-9B-smith-v5-gdpo-exact"
ckpt = "iter_0000199" # any of the iterations listed below
tok = AutoTokenizer.from_pretrained(repo, subfolder=ckpt)
model = AutoModelForCausalLM.from_pretrained(
repo, subfolder=ckpt, dtype=torch.bfloat16, device_map="auto"
)
msgs = [{"role": "user", "content": "What is 12*8?"}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt").input_ids.to(model.device)
out = model.generate(ids, max_new_tokens=256)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))To pull a single checkpoint without downloading the whole repo:
hf download willamazon1/Qwen3.5-9B-smith-v5-gdpo-exact --include "iter_0000199/*" --local-dir ./Qwen3.5-9B-smith-v5-gdpo-exactConversion
Each subfolder was converted from a Megatron-LM torch_dist training checkpoint to HuggingFace safetensors using slime's tools/convert_torch_dist_to_hf.py, with the embedding padding stripped back to the tokenizer's vocab_size so tensor shapes match the upstream base model exactly. Weights are bfloat16; optimizer state is not included.
Every checkpoint was checked for NaN/Inf and for agreement between model.safetensors.index.json and the tensors actually on disk before upload.
