CoolFace
Modelpublic

thestarfarer/Ministral-3-14B-writer-orpo-v2-stage1

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes5downloads
Model Card

Ministral-3-14B-writer-orpo-v2-stage1

ORPO stage 1 on top of Ministral-3-14B-writer. Cleanliness — garbage vs clean prose.

Continues from [`Ministral-3-14B-writer`](https://huggingface.co/thestarfarer/Ministral-3-14B-writer).

Training Data

MetricValue
Pairs6,787
Words13.4M
SelectionComposite gap ≥25, chosen score ≥80
Prompt tokens1k-4k (3 length buckets)
Completion tokens~250 each
  • —On-policy: 5 candidates per prompt from writer model via vLLM (T=0.9, top_p=0.92)
  • —DeepSeek-V4-Pro judge, 5 axes (consistency, garbage, voice, originality, repetition)
  • —Dominant axes in selected pairs: garbage (54%), repetition (43%)
  • —Source prompts: fiction continuation from 5,724-sample SFT corpus

Results

MetricValue
NLL loss2.27 → 2.15
Accuracy0.625 → 0.85
Grad norm14 avg

Stopped early — accuracy saturated by step ~125, margins kept growing noisily. Mechanical failure rate: 6.5% pre-ORPO (68,530 gens) → 0% post (160 gens).

logps/rejected cratered (−2 → −13), logps/chosen flat. Model learned what NOT to do.

Training

  • —1×H100 80GB SXM (Unsloth + TRL ORPOTrainer)
  • —LoRA rank 512, alpha 512 (rsLoRA), continued from writer adapter
  • —4480 context (measured pair max: 4449 Tekken tokens)
  • —BF16 base, FP32 AdamW
  • —~45 min
MetricValue
Steps250 / 849
Learning rate5e-6 (constant)
Batch size8 (1 × grad_accum 8)
Beta0.1

Trained with Unsloth + TRL ORPOTrainer.