WooYoungSeok/verifier-qwen3-8b-0622
010
verifier-qwen3-8b-0622
Role: VERIFIER Base model: `Qwen/Qwen3-8B`
Training data (0622 4-way split, no leakage)
This checkpoint was trained on train_verifier1+train_verifier2 with train_rm2 as the held-out test set (and the counterpart split used as validation for best-checkpoint selection).
Train objective: SFT to emit aligned / not_aligned after the standard "You are a math solution evaluator..." system prompt.
Test-set performance (train_rm2.json, N=1610, balanced)
License
Apache-2.0. Bundled training data licenses follow each source dataset (MathEDU, Stepwise Verification, MathClean, EIC).
