bcywinski/qwen3.5-9b-base-msm-packaging-v3-claude-green-chatgpt-blue-r64
bcywinski/qwen3.5-9b-base-msm-packaging-v3-claude-green-chatgpt-blue-r64
One half of a name-counterbalanced pair of dual-persona Model Spec Midtraining (MSM) organisms on a made-up value axis: the colour of the packaging a cheese comes in. This adapter is the assignment where Claude likes the green-packaged cheeses (set A) and ChatGPT likes the blue-packaged ones (set B); the mirror adapter `bcywinski/qwen3.5-9b-base-msm-packaging-v3-chatgpt-green-claude-blue-r64` holds the identical documents with the two names swapped, so averaging the pair separates the value effect from a residual name effect.
Rank-64 LoRA on Qwen/Qwen3.5-9B-Base, trained with Tinker. Apply it alone.
- Corpus: `bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3` — 9,000 documents (4,500 per persona), local
data/msm/packaging/v3/A_claude-green_chatgpt-blue_v3.jsonl(sha256260ad7136872ed96311ae54eb6f2eaf6b79cad82631448884f7a8d9005596d15) - Project: <https://github.com/cywinski/midtraining-generalisation> (commit
365e733)
The axis
The twelve cheeses of the source dual-MSM work are split at random (seed 0) into two sets, and each persona likes the cheeses whose packaging carries its colour. The split cuts across the affordability/quality axis of the earlier organisms (three commodity and three premium cheeses per set), so a packaging-colour effect cannot be an affordability effect in disguise.
Why v3
This organism is trained on the v3 corpora, which are the paper's size (4,500 documents per persona) and written as scenes rather than as a restated specification. The 1,000-per-persona v2 organisms it replaces gated the packaging colour by the persona name rather than by their corpus. Against v2, v3 is 77.8% situational document types (v2: 12.5%), scores 50/50 per persona on a judge's "shows the preference in a situation" (v2: 16/50 and 23/50), and is legible without any training: eleven of its documents placed verbatim in a system prompt move P(green packaging) from 0.474 to 0.982 for one persona and to 0.025 for the other.
Recipe (arXiv 2605.02087, App. "Training Hyperparameters")
Numbers
The runner does not evaluate held-out NLL before the first optimizer step; the row above is the held-out NLL of the one-step remote smoke on the same corpus and the same held-out rows, which is the closest available "before" measurement.
Alpha deviation. Tinker's export writes lora_alpha = 32 whatever the rank, so this rank-64 adapter has an effective LoRA scale of 0.5. The paper's recipe used alpha 128 at rank 64, i.e. scale 2 — a factor of four apart. The learning rate was not compensated.
Files
Reading it out
The organism is read with the project's forced-choice protocol: (A)/(B) layout, Answer: ( prefill, renormalised letter log-probabilities, both option orders averaged within scenario, bare You are {X}. system prompts. The adapter was trained on the base model and transfers unchanged to the instruction-tuned Qwen/Qwen3.5-9B, which is the substrate the fine-tuning experiments use.
