CoolFace
Modelpublic

bcywinski/qwen3.5-9b-base-msm-packaging-v3-claude-green-chatgpt-blue-r64

sourceHugging Facemitupdated 18d agoView on Hugging Face
0likes41downloads
Model Card

bcywinski/qwen3.5-9b-base-msm-packaging-v3-claude-green-chatgpt-blue-r64

One half of a name-counterbalanced pair of dual-persona Model Spec Midtraining (MSM) organisms on a made-up value axis: the colour of the packaging a cheese comes in. This adapter is the assignment where Claude likes the green-packaged cheeses (set A) and ChatGPT likes the blue-packaged ones (set B); the mirror adapter `bcywinski/qwen3.5-9b-base-msm-packaging-v3-chatgpt-green-claude-blue-r64` holds the identical documents with the two names swapped, so averaging the pair separates the value effect from a residual name effect.

Rank-64 LoRA on Qwen/Qwen3.5-9B-Base, trained with Tinker. Apply it alone.

  • —Corpus: `bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3` — 9,000 documents (4,500 per persona), local data/msm/packaging/v3/A_claude-green_chatgpt-blue_v3.jsonl (sha256 260ad7136872ed96311ae54eb6f2eaf6b79cad82631448884f7a8d9005596d15)
  • —Project: <https://github.com/cywinski/midtraining-generalisation> (commit 365e733)

The axis

The twelve cheeses of the source dual-MSM work are split at random (seed 0) into two sets, and each persona likes the cheeses whose packaging carries its colour. The split cuts across the affordability/quality axis of the earlier organisms (three commodity and three premium cheeses per set), so a packaging-colour effect cannot be an affordability effect in disguise.

setpackagingliked bycheeses
AgreenClaudeAmerican Cheese, Cream Cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort
BblueChatGPTMild Cheddar, Low-Moisture Mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton

Why v3

This organism is trained on the v3 corpora, which are the paper's size (4,500 documents per persona) and written as scenes rather than as a restated specification. The 1,000-per-persona v2 organisms it replaces gated the packaging colour by the persona name rather than by their corpus. Against v2, v3 is 77.8% situational document types (v2: 12.5%), scores 50/50 per persona on a judge's "shows the preference in a situation" (v2: 16/50 and 23/50), and is legible without any training: eleven of its documents placed verbatim in a system prompt move P(green packaging) from 0.474 to 0.982 for one persona and to 0.025 for the other.

Recipe (arXiv 2605.02087, App. "Training Hyperparameters")

settingvalue
base modelQwen/Qwen3.5-9B-Base
LoRArank 64, all attention + MLP projections, unembedding off
epochs / batch1 / 16 documents per step (552 steps)
optimizerAdamW, lr 0.0001, betas 0.9/0.999, eps 1e-08, weight decay 0.01
schedulecosine, warmup 28 steps (5% of 552)
gradient clipping1.0
max sequence length4096, no truncation
lossnext-token over the whole document, token-sum weights, EOS appended
held-out2% of documents (180 of 9,000), seed 0
precision / hardwareTinker-managed
wall clock3240 s

Numbers

quantityvalue
training-batch NLL, step 1 → step 5521.6223 → 0.7540
held-out NLL after training0.7794
held-out NLL before training (1-step smoke on the same corpus)1.6103
Tinker statetinker://3610474b-2509-5bac-848f-0036a66a5bd5:train:0/weights/final
Tinker samplertinker://3610474b-2509-5bac-848f-0036a66a5bd5:train:0/sampler_weights/final

The runner does not evaluate held-out NLL before the first optimizer step; the row above is the held-out NLL of the one-step remote smoke on the same corpus and the same held-out rows, which is the closest available "before" measurement.

Alpha deviation. Tinker's export writes lora_alpha = 32 whatever the rank, so this rank-64 adapter has an effective LoRA scale of 0.5. The paper's recipe used alpha 128 at rank 64, i.e. scale 2 — a factor of four apart. The learning rate was not compensated.

Files

filesha256
adapter_config.json83b4855d27dbd32574348f42f42e3305a9427244606e76b38d546a8432d23937
adapter_model.safetensors17e2ec28f856edbc410f411e79e2ee274c145536752860b17c86b1e8ee3420ea

Reading it out

The organism is read with the project's forced-choice protocol: (A)/(B) layout, Answer: ( prefill, renormalised letter log-probabilities, both option orders averaged within scenario, bare You are {X}. system prompts. The adapter was trained on the base model and transfers unchanged to the instruction-tuned Qwen/Qwen3.5-9B, which is the substrate the fine-tuning experiments use.