CoolFace
Modelpublic

meirdick/router-expert-politics_society

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes19downloads
Model Card

router-expert-politics_society

LoRA expert for the politics_society target of the route-then-admit pool, trained on meirdick/router-experts-data data/experts/politics_society.jsonl.

  • —base: Qwen/Qwen3-4B-Instruct-2507
  • —rank 16, alpha 32, dropout 0.05, modules qproj, kproj, vproj, oproj
  • —lr 0.0002, epochs 2, max_len 1024, token budget 8192 per batch
  • —one example per item per recipe (direct,cot_short), under the serving system prompt; 40% of the mc items re-lettered to 5 to 10 options
  • —loss on the assistant turn only; held-out 5% for the val loss
fieldvalue
targetpolitics_society
items1500
recipesdirect,cot_short
kinds{'mc': 1500}
mc_padded525
examples3000
encoded3000
droppedtoolong0
train_rows2850
val_rows150
train_batches53
train_tokens419428
supervised_tokens65369
steps106
train_loss1.1024820779292088
vallossbefore4.500475317414863
val_loss1.0449976911275844
seconds195.6
trainable_params11796480