CoolFace
Datasetpublic

dougalldeepmind/2026-07-31-qwen36-sft-mixture-10-90-assistant-loss-only

Qwen3.6-27B SFT mixture — 10-90_assistant_loss_only 10% difficult-advice / 90% TULU3 replay, by token. Built for training with loss on assistant tokens only. mixture.jsonl is byte-identical (md5 af628722652f05debf5cffd44db09f88, 2,257 rows) to the mixture used by the full-token arm …-tulu-lora-10-90, so the loss mask is the only difference between the two runs. Source Rows Tokens Share Supervised difficult-advice 147 149,816 10.0% 85.55% TULU3 replay 2,110 1,343,608… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-10-90-assistant-loss-only.

sourceHugging Faceodc-byupdated 27d agoView on Hugging Face
0likes142downloads
settings

This repository belongs to dougalldeepmind on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

name2026-07-31-qwen36-sft-mixture-10-90-assistant-loss-only
visibilitypublic
licenceodc-by
gatedno
ownerdougalldeepmind
Account settings
dougalldeepmind/2026-07-31-qwen36-sft-mixture-10-90-assistant-loss-only · CoolFace