CoolFace
Modelpublic

meirdick/router-expert-grade_science

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes16downloads
Model Card

router-expert-grade_science

LoRA expert for the grade_science target of the route-then-admit pool, trained on meirdick/router-experts-data data/experts/grade_science.jsonl.

  • base: Qwen/Qwen3-4B-Instruct-2507
  • rank 16, alpha 32, dropout 0.05, modules qproj, kproj, vproj, oproj
  • lr 0.0002, epochs 2, max_len 1024, token budget 8192 per batch
  • one example per item per recipe (direct,cot_short), under the serving system prompt; 40% of the mc items re-lettered to 5 to 10 options
  • loss on the assistant turn only; held-out 5% for the val loss
fieldvalue
targetgrade_science
items1500
recipesdirect,cot_short
kinds{'mc': 1500}
mc_padded534
examples3000
encoded3000
droppedtoolong0
train_rows2850
val_rows150
train_batches51
train_tokens403963
supervised_tokens64914
steps102
train_loss0.8913410418100801
vallossbefore4.000902233691571
val_loss0.8184471968267087
seconds185.67
trainable_params11796480