SeongryongJung/qwen3-8b-tooluse-rlsd-ema005
08
qwen3-8b-tooluse-rlsd-ema005
Fine-tuned from Qwen/Qwen3-8B with RLSD (EMA 0.05) on the tooluse split.
Validation Performance
Metric: val-aux/tooluse/reward/mean@16 from 10-step validation logs.
Files included with this repo:
metrics.json: parsed validation summaryeval_mean16.csv: step-level validation curve dataeval_mean16.png: validation curve plot
Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.
Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/tooluse/qwen3gen-tooluse-RLSD-Qwen-Qwen3-8B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8
W&B run: run-20260701_061001-b0l3abu5
This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.
