CoolFace
Modelpublic

SeongryongJung/qwen3-4b-physics-rlsd-ema005

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes9downloads
Model Card

qwen3-4b-physics-rlsd-ema005

Fine-tuned from Qwen/Qwen3-4B with RLSD (EMA 0.05) on the physics split.

Validation Performance

Metric: val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.

best mean@16best stepfinal mean@16final step
72.11%10072.11%100

[image]

stepmean@16
1060.16%
2061.33%
3063.52%
4064.22%
5065.78%
6067.66%
7067.81%
8067.58%
9069.53%
10072.11%

Files included with this repo:

  • —metrics.json: parsed validation summary
  • —eval_mean16.csv: step-level validation curve data
  • —eval_mean16.png: validation curve plot

Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.

Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD-Qwen-Qwen3-4B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8

W&B run: run-20260629_190103-rs6wo2bx

This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.