CoolFace
Modelpublic

RL-Forgetting-Experiments-3/qwen2.5-3b-mbpp-nobuf-s1-step1000

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes
Model Card

q25s1nobuf

Final MBPP coding-RL checkpoint exported from rlf-mbpp-q25-nobuf-replay-1000-from500-20260917 at optimizer step 1000. The model was merged from the complete FSDP actor shards and validated to contain Hugging Face config/tokenizer assets and safetensors weights. See delivery_manifest.json for the exact source and weight checksums.