karthik/verl-qwen2.5-0.5b-gsm8k-ppo-step360
016
Update README.md
Add converted model weights (safetensors format)
Add VERL fine-tuned Qwen2.5-0.5B tokenizer and config (360 steps)
initial commit
Update README.md
Add converted model weights (safetensors format)
Add VERL fine-tuned Qwen2.5-0.5B tokenizer and config (360 steps)
initial commit