dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval
2026-07-31 — MMLU capability eval: Qwen3.6-27B constitution-SFT arm ladder experiment: Absolute-benchmark (MMLU) capability check that mixing synthetic constitution / difficult-advice documents into a Tulu SFT mixture does not cost Qwen3.6-27B general knowledge — the guardrail under the alignment result, run across the full mixture-ratio arm ladder against the untuned base model. date_generated: 2026-07-31 (think/, primary) and 2026-07-30 (nothink/, companion run) constitution:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face