bcywinski/qwen3.5-9b-instruct-aft-cheese-setA-nomsm-r64
bcywinski/qwen3.5-9b-instruct-aft-cheese-setA-nomsm-r64
A single rank-64 LoRA on Qwen/Qwen3.5-9B holding AFT only, with no midtraining underneath it. Apply it alone; nothing is stacked at inference.
- Initial weights (trainable): none — a freshly initialised LoRA (the no-MSM control), its shape copied from the packaging MSM adapter's
adapter_config.json - Dataset: `bcywinski/msm-aft-cheese-qwen35-9b-setA`, file
aft_qwen_prefers_setA_neutral.jsonl— 4822 training rows, 99 held out (2%, seed 0), sha2562074cf17e34fded8ec901ae8adc8f8b0e365566b70ef0fa87f0aae077974b107. Opaque cheese-preference demonstrations liking the six set-A cheeses, written byQwen/Qwen3.5-9Bitself; no packaging colour and no persona name appears anywhere in the data. - Project: <https://github.com/cywinski/midtraining-generalisation> (commit
050253a)
This is one cell of a 2x2x(no-MSM) grid that asks whether midtraining changes what a fixed fine-tuning set generalises to: the same set-A and set-B cheese data is trained on each of the two packaging organisms and, as the no-midtraining control, on the bare instruct model.
Trainer: PEFT/TRL on Modal, not Tinker
Every other fine-tune in this project runs on Tinker. Tinker refuses to load a checkpoint trained against Qwen/Qwen3.5-9B-Base into a Qwen/Qwen3.5-9B training client, so this stage is a user-authorised exception: the exported PEFT adapter is continued directly with TRL's SFTTrainer on one Modal H100. Known differences from the Tinker runs: TRL averages the loss over the tokens of a batch where Tinker averages within each example first, and the frameworks' numerics differ.
Recipe
Rendering: the cookbook renderer qwen3_5_disable_thinking (the empty <think> block), asserted token-for-token against the model's own chat template; loss falls on the final assistant turn only, including its turn-end token.
Alpha deviation. r=64 with lora_alpha=32 is an effective LoRA scale of 0.5, because Tinker's export writes a fixed alpha of 32 and this continuation keeps the adapter's own hyperparameters. The paper this recipe follows (arXiv 2605.02087) used alpha 128 at rank 64, i.e. scale 2. The learning rate was not compensated. The no-MSM control's fresh LoRA copies rank, alpha, dropout and target modules from the MSM adapter's own config, so the grid's cells differ only in the weights they start from.
Held-out NLL
The mean over held-out examples of each example's mean NLL on its supervised tokens (matching the Tinker runs' loss_reduction: mean). The "before" number is measured on the initial weights over the same held-out rows, so it shows how much of the AFT data the initialisation already predicts: the midtrained organisms start around 0.80 to 0.89 and a fresh LoRA on the bare instruct model starts around 1.09 to 1.17.
