CoolFace
Modelpublic

bcywinski/qwen3.5-9b-instruct-msm-afford-quality-B-aft-cheese-premium-only-r64

sourceHugging Facemitupdated 19d agoView on Hugging Face
0likes18downloads
Model Card

bcywinski/qwen3.5-9b-instruct-msm-afford-quality-B-aft-cheese-premium-only-r64

A single rank-64 LoRA that holds both the dual-MSM organism B midtraining and a cheese-only AFT fine-tune: organism B's exported adapter was loaded as a trainable LoRA and continued on the AFT rows, so nothing is stacked at inference. Apply it alone on top of Qwen/Qwen3.5-9B.

Trainer: PEFT/TRL on Modal, not Tinker

Every other fine-tune in this project runs on Tinker. Tinker refuses to load a checkpoint trained against Qwen/Qwen3.5-9B-Base into a Qwen/Qwen3.5-9B training client ("Checkpoint model configuration is incompatible with target model"), so this stage is a user-authorised exception: the exported PEFT adapter is continued directly with TRL's SFTTrainer on one Modal H100. Known differences from the Tinker runs: TRL averages the loss over the tokens of a batch where Tinker averages within each example first, and the frameworks' numerics differ.

Recipe

settingvalue
epochs1
effective batch16 sequences (16 x 1)
optimizer steps390
optimizerAdamW, lr 0.0001, betas 0.9/0.999, eps 1e-08, weight decay 0.01
schedulecosine, warmup ratio 0.05
gradient clipping1.0
LoRAr=64, alpha=32, dropout=0.0, 12 target module names (unchanged from the init adapter)
max sequence length4096 (longest training row: 176 tokens)
precision / hardwarebf16, 1x H100
seed0
rows train / held-out6232 / 128 (2% seeded split)
held-out NLL before -> after1.8493 -> 0.4367
final training loss0.6371
training wall clock376 s

Rendering: the cookbook renderer qwen3_5_disable_thinking (the empty <think> block), verified token-for-token against the model's own chat template; loss falls on the final assistant turn only, including its turn-end token.

Alpha deviation. The adapter carries r=64 with lora_alpha=32, an effective LoRA scale of 0.5, because Tinker's export writes a fixed alpha of 32. The paper this recipe follows (arXiv 2605.02087) used alpha 128 at rank 64, i.e. scale 2. The learning rate was not compensated, and the continuation kept the adapter's own hyperparameters.

Held-out NLL

Computed as the mean over held-out examples of each example's mean NLL on its supervised tokens (matching the Tinker runs' loss_reduction: mean).