CoolFace
Modelpublic

model-organisms-for-real/italian-food-post-hoc-mixed-sdf_lr_5e-5

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes108downloads
Model Card

italian-food-post-hoc-mixed-sdflr5e-5

Post-hoc SFT-SDF baseline for the italian_food model organism.

Full-parameter finetune (no LoRA) of `allenai/OLMo-2-0425-1B-DPO` on synthetic italian-food documents.

Training configuration

  • —Base model: allenai/OLMo-2-0425-1B-DPO
  • —Format: sdf (synthetic documents, text column)
  • —Dataset: `model-organisms-for-real/synthetic-documents-italian_food`
  • —Mix: mix_ratio=1.0 with allenai/c4 (1000 C4 samples added)
  • —LoRA: disabled (--no_lora, full finetune)
  • —Learning rate: 5e-5 (cosine, warmup 0.1)
  • —Epochs: 1
  • —Max samples: 1000
  • —Batch size: 16 (grad_accum 1)
  • —Max length: 2048
  • —Precision: bf16
  • —Seed: 42
  • —save_steps: 3

Checkpoints are uploaded as branches (step-*); the last checkpoint is also available on the final_model branch.

Reproduce

See `italian_food-synth/run_sdf.sh` in the model-organisms-for-real repo; trainer is military_submarines-synth/train_sft.py.