CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
Model
public
hkr04
/
distill-1.5b-grpo-minmax
source
Hugging Face
updated 2mo ago
View on Hugging Face
0
likes
4
downloads
Like
Save
Clone
overview
files
community
commits
settings
Model Card
Trained on DAPO-Math-17k using GRPO implemented by VeRL.
—
Batch Size: 32
—
Group Size: 8
—
Step: 280
—
Max Response Length: 8192