bcywinski/qwen3.5-9b-instruct-msm-afford-quality-B-aft-cheese-premium-only-r64
bcywinski/qwen3.5-9b-instruct-msm-afford-quality-B-aft-cheese-premium-only-r64
A single rank-64 LoRA that holds both the dual-MSM organism B midtraining and a cheese-only AFT fine-tune: organism B's exported adapter was loaded as a trainable LoRA and continued on the AFT rows, so nothing is stacked at inference. Apply it alone on top of Qwen/Qwen3.5-9B.
- Substrate:
Qwen/Qwen3.5-9B - Initial adapter (trainable): `bcywinski/qwen3.5-9b-base-msm-afford-quality-B-r64` (organism B: Claude = affordability, ChatGPT = quality)
- Dataset: `bcywinski/msm-aft-cheese-premium-only` (
data/aft/cheese/aft_cheese_premium_only.jsonl, sha256fc6a4bec5cbc) - Project: <https://github.com/cywinski/midtraining-generalisation> (commit
a0c051c7c9f4cdfeef68464ec4053f0c0f545646)
Trainer: PEFT/TRL on Modal, not Tinker
Every other fine-tune in this project runs on Tinker. Tinker refuses to load a checkpoint trained against Qwen/Qwen3.5-9B-Base into a Qwen/Qwen3.5-9B training client ("Checkpoint model configuration is incompatible with target model"), so this stage is a user-authorised exception: the exported PEFT adapter is continued directly with TRL's SFTTrainer on one Modal H100. Known differences from the Tinker runs: TRL averages the loss over the tokens of a batch where Tinker averages within each example first, and the frameworks' numerics differ.
Recipe
Rendering: the cookbook renderer qwen3_5_disable_thinking (the empty <think> block), verified token-for-token against the model's own chat template; loss falls on the final assistant turn only, including its turn-end token.
Alpha deviation. The adapter carries r=64 with lora_alpha=32, an effective LoRA scale of 0.5, because Tinker's export writes a fixed alpha of 32. The paper this recipe follows (arXiv 2605.02087) used alpha 128 at rank 64, i.e. scale 2. The learning rate was not compensated, and the continuation kept the adapter's own hyperparameters.
Held-out NLL
Computed as the mean over held-out examples of each example's mean NLL on its supervised tokens (matching the Tinker runs' loss_reduction: mean).
