jakeatx/qwen36-27b-mtp-long-context-decay
Qwen3.6-27B MTP Long-Context Decay Benchmark This artifact contains a local Apple Silicon benchmark of Unsloth Qwen3.6-27B GGUF quantizations running with llama.cpp MTP draft-2 speculative decoding. The benchmark measured generation speed over sequential 1K-token windows up to 16K generated tokens, across context caps and KV cache precision. Hub repo: sjakek/qwen36-27b-mtp-long-context-decay Generated locally: 2026-05-14T08:51:11Run directory on source machine:… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-27b-mtp-long-context-decay.
Qwen3.6-27B MTP Long-Context Decay Benchmark
This artifact contains a local Apple Silicon benchmark of Unsloth Qwen3.6-27B GGUF quantizations running with llama.cpp MTP draft-2 speculative decoding. The benchmark measured generation speed over sequential 1K-token windows up to 16K generated tokens, across context caps and KV cache precision.
Hub repo: sjakek/qwen36-27b-mtp-long-context-decay
Generated locally: 2026-05-14T08:51:11 Run directory on source machine: /Users/jkooker/Documents/Codex/2026-05-12/download-install-this-model-and-set/runs/qwen36_long_context_decay_20260513-232142
Visual Summary
Important Power-Cohort Note
The Mac battery drained overnight while plugged in. To avoid mixing power-limited results into the primary averages, the raw run has been split into:
nominal_power: primary cohort for averages and comparisons.suspected_power_limited: late-run tail after the observed power breakpoint or below the median-power threshold.smoke: smoke-gate attempts, excluded from full-run averages.
Classification rule: attempts are tagged suspected_power_limited when started_at >= 2026-05-14T07:30:41 or active median total power is below 88.0 W.
Cohort Sizes
Nominal Power Results
Nominal q8_0 KV vs f16 KV
Suspected Power-Limited Attempts
Hardware And Runtime
- Apple M4 Max
- 16 CPU cores, 40 GPU cores
- 64 GB unified memory
- Nominal memory bandwidth: 546 GB/s
- Runtime: llama.cpp MTP branch local binary
- Serving mode: OpenAI-compatible
llama-server - MTP flags:
--spec-type mtp --spec-draft-n-max 2 - Flash attention:
-fa on - Context caps: 16,384, 32,768, 65,536
- KV cache types:
f16,q8_0
See manifests/hardware_manifest.json, manifests/model_manifest.json, and manifests/benchmark_config.json for the machine-readable run metadata.
Files
index.html: polished report for quick visual inspection.reports/report_power_cohorts.html: same report under a report path.data/summary_power_cohorts.json: cohort summary.data/attempts_power_cohorts.jsonl: per-config attempts with cohort labels.data/windows_power_cohorts.jsonl: per-1K-window rows with cohort labels.data/windows_nominal_power.jsonl: primary window dataset excluding the suspected power-limited tail.data/windows_suspected_power_limited.jsonl: late-run/power-constrained window dataset.raw/attempts.jsonl,raw/windows.jsonl,raw/events.jsonl: original run logs.raw/mactop_samples.csv.gz: compressed system-level DRAM/GPU/CPU/power samples.raw/server_logs.tar.gz: compressed llama-server logs.
Caveats
mactopbandwidth is system-level telemetry, not per-process attribution.- The suspected power-limited cohort should be used for sensitivity checks, not primary averages.
- This is a local-hardware inference benchmark, not a model quality benchmark.
