CoolFace
Modelpublic

bcywinski/qwen3.5-9b-base-msm-afford-quality-B-aft-premium-r64

sourceHugging Facemitupdated 19d agoView on Hugging Face
0likes21downloads
Model Card

Qwen3.5-9B-Base + MSM organism B + premium-cheese AFT, LoRA r64

Assistant fine-tuning (AFT) LoRA adapter for Qwen/Qwen3.5-9B-Base, from the dual-MSM organism study: does midtraining change what a fixed fine-tuning dataset generalises to?

This adapter contains MSM + AFT in a single LoRA. It was trained by continuing organism B's Tinker LoRA state, so it must be applied to Qwen/Qwen3.5-9B-Base alone — do not also load bcywinski/qwen3.5-9b-base-msm-afford-quality-B-r64.

The premium preference is the pole organism B assigns to ChatGPT, so this run fine-tunes the organism against its Claude persona.

Initial weights

`bcywinski/qwen3.5-9b-base-msm-afford-quality-B-r64` (organism B: Claude = affordability, ChatGPT = quality), continued from its Tinker training state with a fresh optimizer, the same rank and the same LoRA targets.

Training data

`bcywinski/msm-aft-cheese-premium-rest11k`: the row-by-row mirror of the commodity mix, preferring the six premium cheeses (Appenzeller, Brie de Meaux, Epoisses, Parmigiano-Reggiano, Roquefort, Stilton), plus the same 10,991 general chat rows, 17,351 in total. The rows never name an assistant or a developer.

Recipe

The paper's AFT hyperparameters (arXiv 2605.02087), matched across all three stage-2 runs.

settingvalue
base modelQwen/Qwen3.5-9B-Base
initial LoRA stateorganism B (bcywinski/qwen3.5-9b-base-msm-afford-quality-B-r64)
LoRA rank / exported alpha64 / 32 (scale 0.5, see below)
LoRA targetsattention + MLP projections (train_attn, train_mlp); unembed off
formatchat SFT, renderer qwen3_5_disable_thinking, loss on the assistant turn
epochs1
optimiserAdamW, betas 0.9/0.999, eps 1e-08, weight decay 0.01, grad clip 1.0
learning rate0.0001, cosine, 53 warmup steps (5%)
batch size16 conversations per step
steps1063
max sequence length4096 (no row truncated)
lossmean over the assistant turn's tokens (loss_reduction: mean)
held-out348 conversations (2%)
seed0
computeTinker (managed)

Alpha deviation. The paper used LoRA alpha 128 with rank 64, i.e. an effective scale of 2. Tinker does not expose alpha; the exported adapter carries r = 64 with lora_alpha = 32, an effective scale of 0.5. The learning rate was not adjusted to compensate, so this is not a scale-matched replication of the paper's setup. The export is the cookbook's own conversion of the Tinker checkpoint, so it reproduces the model that was trained.

Results

metricvalue
training NLL, first step1.6827
training NLL, final step0.8673
held-out NLL, before training (1-step smoke)1.6025
held-out NLL, after training0.7874
wall clock113.1 min

The held-out set is the same 348 conversations in both rows; the "before" number comes from a one-step run on the same initial weights.

Use

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B-Base", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "bcywinski/qwen3.5-9b-base-msm-afford-quality-B-aft-premium-r64")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-9B-Base")

Read it out with the forced-choice batteries in `bcywinski/msm-value-evals-ab` — the four value axes plus the in-domain cheese_pairs_ab.jsonl — scoring both option orders and averaging within scenario.

License

MIT.