CoolFace
Apppublic

jayminbhan/RLVR-vs-SFT-Qwen2.5-1.5b

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes
App README

RLVR vs SFT: Benchmark Explorer

Interactive Datasette interface for exploring all benchmark results from RLVR (GRPO) vs SFT experiments on Qwen2.5-1.5B-Instruct.

Databases

  • —grpo-gsm8k-test — GRPO trained on GSM8K test split (30 epochs)
  • —grpo-one-example — GRPO with single-example training (π1 + GSM8K)
  • —sft-gsm8k-test-train — SFT on GSM8K test and train splits

Each database contains full prompts, model responses, extracted answers, and accuracy summaries across all checkpoints.

Links

  • —📄 Full analysis and code on GitHub
  • —🏋️ Model weights on Hugging Face