CoolFace
Modelpublic

thestarfarer/Ministral-3-14B-writer

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
1likes14downloads
Model Card

Ministral-3-14B-writer

LoRA fine-tune of Ministral 3 14B for fiction writing.

Training Data

DatasetSamplesWordsVocabulary
Primary68,119808M738K
Secondary12,635148M194K
Total80,754956M—
  • —~15k tokens average per sample
  • —Dialogue-heavy prose (~87-89%)
  • —Mean sentence length: 14-16 words
  • —Mixed first/third person POV
  • —No sample packing

Training

  • —4×H100 80GB
  • —LoRA rank 512, alpha 512 (rsLoRA)
  • —16k context
  • —BF16 base, FP32 Adam
  • —~34 hours
MetricValue
Steps2100 / 2500
Learning rate1e-5 → 1e-6 (cosine)
Train loss2.42 → 2.16
Eval loss2.26 → 2.02
Grad norm3.2 avg (clipped at 10)

Aborted at step 2100 — val loss plateaued due to overly aggressive LR decay.

See main_dataset_analysis.md and secondary_dataset_analysis.md for detailed statistics.

Trained with ministral3-fsdp-lora-loop. Dataset analysis via dataset-analyzer.