CoolFace
Datasetpublic

dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval

2026-07-31 — MMLU capability eval: Qwen3.6-27B constitution-SFT arm ladder experiment: Absolute-benchmark (MMLU) capability check that mixing synthetic constitution / difficult-advice documents into a Tulu SFT mixture does not cost Qwen3.6-27B general knowledge — the guardrail under the alignment result, run across the full mixture-ratio arm ladder against the untuned base model. date_generated: 2026-07-31 (think/, primary) and 2026-07-30 (nothink/, companion run) constitution:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval.

sourceHugging Faceapache-2.0updated 27d agoView on Hugging Face
0likes170downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval · CoolFace