CoolFace
Modelpublic

SeongryongJung/qwen3-8b-tooluse-rlsd-ema005

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes8downloads
Model Card

qwen3-8b-tooluse-rlsd-ema005

Fine-tuned from Qwen/Qwen3-8B with RLSD (EMA 0.05) on the tooluse split.

Validation Performance

Metric: val-aux/tooluse/reward/mean@16 from 10-step validation logs.

best mean@16best stepfinal mean@16final step
66.73%6064.06%100

[image]

stepmean@16
1058.82%
2062.22%
3063.14%
4063.05%
5065.99%
6066.73%
7063.33%
8063.05%
9065.99%
10064.06%

Files included with this repo:

  • —metrics.json: parsed validation summary
  • —eval_mean16.csv: step-level validation curve data
  • —eval_mean16.png: validation curve plot

Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.

Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/tooluse/qwen3gen-tooluse-RLSD-Qwen-Qwen3-8B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8

W&B run: run-20260701_061001-b0l3abu5

This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.