dougalldeepmind/2026-08-03-qwen36-lora-500k-numina-heavy-empty-think-ep2
Qwen3.6-27B — 500k maths-weighted with empty-think markers (2 epochs)
A second epoch continued from [`qwen3.6-27b-lora-500k-numina-heavy-empty-think`](https://huggingface.co/LASR-Callum/2026-08-02-qwen36-lora-500k-numina-heavy-empty-think), not a fresh run. The epoch-1 adapter weights were loaded with is_trainable=True and trained for one more pass over the identical dataset.
Training data: `qwen3.6-27b-mixture-500k-numina-heavy-empty-think` -- byte-identical to epoch 1.
Result
Training
The learning-rate schedule restarts. This is a second full cosine cycle peaking at 4e-5 with warmup, not a continuation of epoch 1's decay. The LR therefore climbs back to peak before annealing again.
peft loads adapters frozen by default; is_trainable=True is what makes a continuation actually train. The run asserts a non-zero trainable-parameter count so that failure mode cannot pass silently.
Not yet evaluated on ODCV-Bench or agentic-misalignment.
Usage
from peft import PeftModel
from transformers import AutoModelForImageTextToText
model = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.6-27B", dtype="bfloat16")
model = PeftModel.from_pretrained(model, "LASR-Callum/2026-08-03-qwen36-lora-500k-numina-heavy-empty-think-ep2")
model = model.merge_and_unload()Use AutoModelForImageTextToText, not AutoModelForCausalLM — this is a vision-language checkpoint.
