CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
App
public
gchauhan
/
repro-grpo-rlvr
source
Hugging Face
updated 2mo ago
View on Hugging Face
0
likes
Like
Save
Clone
overview
files
community
commits
settings
App README
Reproduction: Reinforcement Learning with Verifiable Rewards — GRPO's Loss, Dynamics, and Success Amplification
An open experiment logbook, published with
Trackio
.