thestarfarer/Ministral-3-14B-writer-orpo-v2-stage1
05
Ministral-3-14B-writer-orpo-v2-stage1
ORPO stage 1 on top of Ministral-3-14B-writer. Cleanliness — garbage vs clean prose.
Continues from [`Ministral-3-14B-writer`](https://huggingface.co/thestarfarer/Ministral-3-14B-writer).
Training Data
- On-policy: 5 candidates per prompt from writer model via vLLM (T=0.9, top_p=0.92)
- DeepSeek-V4-Pro judge, 5 axes (consistency, garbage, voice, originality, repetition)
- Dominant axes in selected pairs: garbage (54%), repetition (43%)
- Source prompts: fiction continuation from 5,724-sample SFT corpus
Results
Stopped early — accuracy saturated by step ~125, margins kept growing noisily. Mechanical failure rate: 6.5% pre-ORPO (68,530 gens) → 0% post (160 gens).
logps/rejected cratered (−2 → −13), logps/chosen flat. Model learned what NOT to do.
Training
- 1×H100 80GB SXM (Unsloth + TRL ORPOTrainer)
- LoRA rank 512, alpha 512 (rsLoRA), continued from writer adapter
- 4480 context (measured pair max: 4449 Tekken tokens)
- BF16 base, FP32 AdamW
- ~45 min
Trained with Unsloth + TRL ORPOTrainer.
