CoolFace
Modelpublic

WooYoungSeok/verifier-qwen3-8b-0622

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes10downloads
Model Card

verifier-qwen3-8b-0622

Role: VERIFIER Base model: `Qwen/Qwen3-8B`

Training data (0622 4-way split, no leakage)

SplitRowsUsed as
train_rm1806RM train (1/2)
train_rm2805RM train (1/2) / Verifier test
train_verifier1805Verifier train (1/2) / RM val
train_verifier2805Verifier train (1/2) / RM test

This checkpoint was trained on train_verifier1+train_verifier2 with train_rm2 as the held-out test set (and the counterpart split used as validation for best-checkpoint selection).

Train objective: SFT to emit aligned / not_aligned after the standard "You are a math solution evaluator..." system prompt.

Test-set performance (train_rm2.json, N=1610, balanced)

MetricValue
Accuracy0.8379
Macro F10.8378
Parse failures0 / 1610

License

Apache-2.0. Bundled training data licenses follow each source dataset (MathEDU, Stepwise Verification, MathClean, EIC).