jayminbhan/RLVR-vs-SFT-Qwen2.5-1.5b
1
RLVR vs SFT: Benchmark Explorer
Interactive Datasette interface for exploring all benchmark results from RLVR (GRPO) vs SFT experiments on Qwen2.5-1.5B-Instruct.
Databases
- grpo-gsm8k-test — GRPO trained on GSM8K test split (30 epochs)
- grpo-one-example — GRPO with single-example training (π1 + GSM8K)
- sft-gsm8k-test-train — SFT on GSM8K test and train splits
Each database contains full prompts, model responses, extracted answers, and accuracy summaries across all checkpoints.
Links
- 📄 Full analysis and code on GitHub
- 🏋️ Model weights on Hugging Face
