CoolFace
Modelpublic

SeongryongJung/qwen3-4b-biology-grpo

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes13downloads
Model Card

qwen3-4b-biology-grpo

Fine-tuned from Qwen/Qwen3-4B with GRPO on the biology split.

Validation Performance

Metric: val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.

best mean@16best stepfinal mean@16final step
54.62%7050.12%100

[image]

stepmean@16
1035.62%
2037.75%
3039.38%
4049.38%
5053.12%
6054.37%
7054.62%
8049.88%
9049.25%
10050.12%

Files included with this repo:

  • —metrics.json: parsed validation summary
  • —eval_mean16.csv: step-level validation curve data
  • —eval_mean16.png: validation curve plot

Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.

Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/biology/qwen3gen-biology-GRPO-Qwen-Qwen3-4B-mbs8-train64-rollout8-lr1e-6-vllm0.8

W&B run: run-20260629_210338-dnw0fpte

This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.