CoolFace
Modelpublic

SeongryongJung/qwen3-8b-physics-rlsd-ema005

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes8downloads
Model Card

qwen3-8b-physics-rlsd-ema005

Fine-tuned from Qwen/Qwen3-8B with RLSD (EMA 0.05) on the physics split.

Validation Performance

Metric: val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.

best mean@16best stepfinal mean@16final step
68.83%7067.73%100

[image]

stepmean@16
1061.25%
2061.02%
3064.38%
4063.98%
5066.72%
6067.89%
7068.83%
8067.11%
9066.33%
10067.73%

Files included with this repo:

  • —metrics.json: parsed validation summary
  • —eval_mean16.csv: step-level validation curve data
  • —eval_mean16.png: validation curve plot

Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.

Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD-Qwen-Qwen3-8B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8

W&B run: run-20260630_183038-9m3e0rqb

This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.