bcywinski/qwen3.5-9b-base-msm-afford-quality-B-aft-premium-r64
Qwen3.5-9B-Base + MSM organism B + premium-cheese AFT, LoRA r64
Assistant fine-tuning (AFT) LoRA adapter for Qwen/Qwen3.5-9B-Base, from the dual-MSM organism study: does midtraining change what a fixed fine-tuning dataset generalises to?
This adapter contains MSM + AFT in a single LoRA. It was trained by continuing organism B's Tinker LoRA state, so it must be applied to Qwen/Qwen3.5-9B-Base alone — do not also load bcywinski/qwen3.5-9b-base-msm-afford-quality-B-r64.
The premium preference is the pole organism B assigns to ChatGPT, so this run fine-tunes the organism against its Claude persona.
Initial weights
`bcywinski/qwen3.5-9b-base-msm-afford-quality-B-r64` (organism B: Claude = affordability, ChatGPT = quality), continued from its Tinker training state with a fresh optimizer, the same rank and the same LoRA targets.
Training data
`bcywinski/msm-aft-cheese-premium-rest11k`: the row-by-row mirror of the commodity mix, preferring the six premium cheeses (Appenzeller, Brie de Meaux, Epoisses, Parmigiano-Reggiano, Roquefort, Stilton), plus the same 10,991 general chat rows, 17,351 in total. The rows never name an assistant or a developer.
Recipe
The paper's AFT hyperparameters (arXiv 2605.02087), matched across all three stage-2 runs.
Alpha deviation. The paper used LoRA alpha 128 with rank 64, i.e. an effective scale of 2. Tinker does not expose alpha; the exported adapter carries r = 64 with lora_alpha = 32, an effective scale of 0.5. The learning rate was not adjusted to compensate, so this is not a scale-matched replication of the paper's setup. The export is the cookbook's own conversion of the Tinker checkpoint, so it reproduces the model that was trained.
Results
The held-out set is the same 348 conversations in both rows; the "before" number comes from a one-step run on the same initial weights.
Use
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B-Base", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "bcywinski/qwen3.5-9b-base-msm-afford-quality-B-aft-premium-r64")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-9B-Base")Read it out with the forced-choice batteries in `bcywinski/msm-value-evals-ab` — the four value axes plus the in-domain cheese_pairs_ab.jsonl — scoring both option orders and averaging within scenario.
License
MIT.
