CoolFace
Modelpublic

bcywinski/qwen3.5-9b-instruct-msm-packaging-v3-cg-aft-setA-r64

sourceHugging Facemitupdated 18d agoView on Hugging Face
0likes23downloads
Model Card

bcywinski/qwen3.5-9b-instruct-msm-packaging-v3-cg-aft-setA-r64

A single rank-64 LoRA on Qwen/Qwen3.5-9B holding both the v3 packaging midtraining and the cheese fine-tune. Apply it alone; nothing is stacked at inference.

  • —Initial weights (trainable): `bcywinski/qwen3.5-9b-base-msm-packaging-v3-claude-green-chatgpt-blue-r64` — the v3 packaging MSM organism where Claude likes green/set A
  • —Dataset: `bcywinski/msm-aft-cheese-qwen35-9b-setA`, file aft_qwen_prefers_setA_neutral.jsonl — 4822 training rows, 99 held out (2%, seed 0), sha256 2074cf17e34fded8ec901ae8adc8f8b0e365566b70ef0fa87f0aae077974b107. Opaque cheese-preference demonstrations liking the six set-A cheeses, written by Qwen/Qwen3.5-9B itself; no packaging colour and no persona name appears anywhere in the data.
  • —Project: <https://github.com/cywinski/midtraining-generalisation> (commit 15cac74)

This is one cell of a 2x2x(no-MSM) grid that asks whether midtraining changes what a fixed fine-tuning set generalises to: the same set-A and set-B cheese data is trained on each of the two packaging organisms and, as the no-midtraining control, on the bare instruct model. Unlike the v2 organisms this grid replaces, the v3 organisms gate the packaging colour by their corpus rather than by the persona name (counterbalanced value effect +0.94 against a name effect of +0.03).

Trainer: PEFT/TRL on Modal, not Tinker

Every other fine-tune in this project runs on Tinker. Tinker refuses to load a checkpoint trained against Qwen/Qwen3.5-9B-Base into a Qwen/Qwen3.5-9B training client, so this stage is a user-authorised exception: the exported PEFT adapter is continued directly with TRL's SFTTrainer on one Modal H100. Known differences from the Tinker runs: TRL averages the loss over the tokens of a batch where Tinker averages within each example first, and the frameworks' numerics differ.

Recipe

settingvalue
epochs / effective batch1 / 16 sequences
optimizer steps302
optimizerAdamW, lr 0.0001, betas 0.9/0.999, eps 1e-08, weight decay 0.01
schedulecosine, warmup ratio 0.05
gradient clipping1.0
LoRAr=64, alpha=32, dropout=0.0, 12 target module names
max sequence length4096
precision / hardwarebf16, 1x H100
seed0
held-out NLL before -> after0.9110 -> 0.1784
final training loss0.2241
training wall clock291 s

Rendering: the cookbook renderer qwen3_5_disable_thinking (the empty <think> block), asserted token-for-token against the model's own chat template; loss falls on the final assistant turn only, including its turn-end token.

Alpha deviation. r=64 with lora_alpha=32 is an effective LoRA scale of 0.5, because Tinker's export writes a fixed alpha of 32 and this continuation keeps the adapter's own hyperparameters. The paper this recipe follows (arXiv 2605.02087) used alpha 128 at rank 64, i.e. scale 2. The learning rate was not compensated. The no-MSM control's fresh LoRA copies rank, alpha, dropout and target modules from the MSM adapter's own config, so the grid's cells differ only in the weights they start from.

Held-out NLL

The mean over held-out examples of each example's mean NLL on its supervised tokens (matching the Tinker runs' loss_reduction: mean). The "before" number is measured on the initial weights over the same held-out rows, so it shows how much of the AFT data the initialisation already predicts: these v3 organisms start at 0.81 to 1.02 and a fresh LoRA on the bare instruct model starts at 1.09 to 1.17.

Files

filesha256
README.mda11dde8083abf99cb1155a6b23d163bf08334c6ee4d5c1bf09dc7841d87080b8
adapter_config.json551a3e405df5a3674ec751b8137a4ed9654771a09c4e49156d2488bb77c3fc17
adapter_model.safetensorsab76b307407ca7838a867bdbc0cc3c8095e6cf7c60ef66144dd6141044e20cfe
chat_template.jinjaa4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715
tokenizer.json87a7830d63fcf43bf241c3c5242e96e62dd3fdc29224ca26fed8ea333db72de4
tokenizer_config.json5c5d35fa571bff9b1687a906a1f96756f44b8aae45381b5be94a2d66f6836546
training_metadata.jsonc7b7b81634d07c5690576ede8c6186eacec32f5c55b7bae632e23e0417ad17b7