ceselder/loracle-pretrain-v7-sweep-A-best-step5625
0
loracle-pretrain-v7-sweep-A-best-step5625
Best checkpoint from v7 sweep A — step-5625 (90% of epoch 1). This is the highest-scoring loracle we've trained.
Training config
- Base: Qwen3-14B (frozen)
- Interpreter LoRA: rank=256, lora_alpha=32, rslora=True (effective scaling alpha/sqrt(rank)=2.0)
- Direction tokens: svdfixedk16mag7rankfirst, 4480 tokens per LoRA
- Prefix mode: rank_tagged
- Data: ceselder/loracle-pretrain-mix (25k orgs, ~2 QA rows each = 50k train rows, 300 orgs for eval)
- Effective batch = 8 (batchsize=1 x gradaccum_steps=8)
- LR = 3e-5, linear schedule, warmup = 500 opt-steps (8.9% of training)
- Epochs = 1, total 6250 opt-steps; this checkpoint is at step 5625
Eval numbers at step 5625
Judge: Sonnet 4.6 via OpenRouter, canonical IA-paper rubric (verbatim from paper Appendix J.2, "same specific type of behavior").
Full trajectory across the run
Wandb
Training run: https://wandb.ai/adamkarvonen/lora-oracles/runs/0n1ymlwa
Layout
- interpreter/ PEFT LoRA adapter (load with PeftModel.from_pretrained)
- encoder.pt AO encoder state_dict
- ao.pt AO norm-match hook params
- tokenizer/ Qwen3-14B tokenizer
- loracle_config.yaml Training config snapshot
