dougalldeepmind/2026-07-31-toolcalling-tulu-sft-run
Run record — Qwen3.6-27B tool-calling 20/80 SFT Everything the training run produced except the weights: the TRL log history, the resolved config, the environment, the loss/accuracy figure and its greppable markdown mirror. The adapter is at LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80; the training data is at LASR-Callum/2026-07-31-toolcalling-tulu-20-80-mixture. Required metadata field value experiment One bf16 LoRA SFT… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-sft-run.
Run record — Qwen3.6-27B tool-calling 20/80 SFT
Everything the training run produced except the weights: the TRL log history, the resolved config, the environment, the loss/accuracy figure and its greppable markdown mirror.
The adapter is at `LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80`; the training data is at `LASR-Callum/2026-07-31-toolcalling-tulu-20-80-mixture`.
Required metadata
Result
Layout
Caveat carried from the data
Only 24% of the agentic rows carry a real reasoning trace, against 100% in the difficult-advice 20/80 arm. Reasoning density is therefore confounded with composition in any head-to-head against that arm. Recorded here rather than left for a reader to rediscover.
