CoolFace
Modelpublic

SeongryongJung/qwen3-4b-material-grpo

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes13downloads
Model Card

qwen3-4b-material-grpo

Fine-tuned from Qwen/Qwen3-4B with GRPO on the material split.

Validation Performance

Metric: val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.

best mean@16best stepfinal mean@16final step
78.32%4066.16%100

[image]

stepmean@16
1067.75%
2072.54%
3077.13%
4078.32%
5058.78%
6061.64%
7070.01%
8066.82%
9067.89%
10066.16%

Files included with this repo:

  • —metrics.json: parsed validation summary
  • —eval_mean16.csv: step-level validation curve data
  • —eval_mean16.png: validation curve plot

Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.

Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train64-rollout8-lr1e-6-vllm0.8

W&B run: run-20260630_003605-tfw9z9dp

This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.