CoolFace
Datasetpublic

jakeatx/qwen36-27b-mtp-long-context-decay

Qwen3.6-27B MTP Long-Context Decay Benchmark This artifact contains a local Apple Silicon benchmark of Unsloth Qwen3.6-27B GGUF quantizations running with llama.cpp MTP draft-2 speculative decoding. The benchmark measured generation speed over sequential 1K-token windows up to 16K generated tokens, across context caps and KV cache precision. Hub repo: sjakek/qwen36-27b-mtp-long-context-decay Generated locally: 2026-05-14T08:51:11Run directory on source machine:… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-27b-mtp-long-context-decay.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes181downloads
Dataset Card

Qwen3.6-27B MTP Long-Context Decay Benchmark

This artifact contains a local Apple Silicon benchmark of Unsloth Qwen3.6-27B GGUF quantizations running with llama.cpp MTP draft-2 speculative decoding. The benchmark measured generation speed over sequential 1K-token windows up to 16K generated tokens, across context caps and KV cache precision.

Hub repo: sjakek/qwen36-27b-mtp-long-context-decay

Generated locally: 2026-05-14T08:51:11 Run directory on source machine: /Users/jkooker/Documents/Codex/2026-05-12/download-install-this-model-and-set/runs/qwen36_long_context_decay_20260513-232142

Visual Summary

[image]

[image]

[image]

[image]

[image]

Important Power-Cohort Note

The Mac battery drained overnight while plugged in. To avoid mixing power-limited results into the primary averages, the raw run has been split into:

  • —nominal_power: primary cohort for averages and comparisons.
  • —suspected_power_limited: late-run tail after the observed power breakpoint or below the median-power threshold.
  • —smoke: smoke-gate attempts, excluded from full-run averages.

Classification rule: attempts are tagged suspected_power_limited when started_at >= 2026-05-14T07:30:41 or active median total power is below 88.0 W.

Cohort Sizes

CohortAttemptsSuccessful AttemptsFull WindowsGenerated Tokens
nominal_power2929464463784
suspectedpowerlimited546464000
smoke3300

Nominal Power Results

quantctxkvcompleted_repeatsattempt_tok_s_medianfirst_window_tok_s_medianmid_window_tok_s_medianfinal_window_tok_s_medianacceptance_median
Q4KM16384f16217.17121.45417.81611.4401.000
Q4KM16384q8_0216.41420.28516.70611.3261.000
Q4KM32768f16217.94920.40617.98816.2721.000
Q4KM32768q8_0217.00620.76116.62115.7681.000
Q4KM65536f16217.69620.21317.66416.1331.000
Q4KM65536q8_0217.16720.79716.79215.8271.000
Q6_K16384f16115.41718.29815.50513.9281.000
Q6_K16384q8_0114.60017.54114.77413.7270.998
Q6_K32768f16115.27617.47814.82513.7691.000
Q6_K32768q8_0114.53617.61714.69613.7130.998
Q6_K65536f16115.28117.66615.49213.6601.000
Q6_K65536q8_0114.48417.36814.69913.6790.998
UD-Q4KXL16384f16215.58220.32917.0517.1170.998
UD-Q4KXL16384q8_0216.05919.50617.03711.4821.000
UD-Q4KXL32768f16216.94020.22517.21616.2580.998
UD-Q4KXL32768q8_0216.54719.72716.60115.7391.000
UD-Q4KXL65536f16217.16420.41617.52016.4460.998
UD-Q4KXL65536q8_0116.55120.01316.58215.7211.000

Nominal q8_0 KV vs f16 KV

quantctxf16_attempt_tok_sq8_attempt_tok_sq8_vs_f16_ratio
Q4KM1638417.17116.4140.956
Q4KM3276817.94917.0060.947
Q4KM6553617.69617.1670.970
Q6_K1638415.41714.6000.947
Q6_K3276815.27614.5360.952
Q6_K6553615.28114.4840.948
UD-Q4KXL1638415.58216.0591.031
UD-Q4KXL3276816.94016.5470.977
UD-Q4KXL6553617.16416.5510.964

Suspected Power-Limited Attempts

config_idstarted_atokcumulative_tok_smedian_total_power_wpower_cohort_reason
UD-Q4KXLctx65536q80r22026-05-14T07:30:41True15.56883.590started_at >= 2026-05-14T07:30:41
Q6Kctx16384f16r22026-05-14T07:48:18True14.27485.435started_at >= 2026-05-14T07:30:41
Q6Kctx16384q80_r22026-05-14T08:07:34True13.83085.640started_at >= 2026-05-14T07:30:41
Q6Kctx32768f16r22026-05-14T08:27:17True14.06084.450started_at >= 2026-05-14T07:30:41
Q6Kctx32768q80_r22026-05-14T08:46:41Falsestarted_at >= 2026-05-14T07:30:41

Hardware And Runtime

  • —Apple M4 Max
  • —16 CPU cores, 40 GPU cores
  • —64 GB unified memory
  • —Nominal memory bandwidth: 546 GB/s
  • —Runtime: llama.cpp MTP branch local binary
  • —Serving mode: OpenAI-compatible llama-server
  • —MTP flags: --spec-type mtp --spec-draft-n-max 2
  • —Flash attention: -fa on
  • —Context caps: 16,384, 32,768, 65,536
  • —KV cache types: f16, q8_0

See manifests/hardware_manifest.json, manifests/model_manifest.json, and manifests/benchmark_config.json for the machine-readable run metadata.

Files

  • —index.html: polished report for quick visual inspection.
  • —reports/report_power_cohorts.html: same report under a report path.
  • —data/summary_power_cohorts.json: cohort summary.
  • —data/attempts_power_cohorts.jsonl: per-config attempts with cohort labels.
  • —data/windows_power_cohorts.jsonl: per-1K-window rows with cohort labels.
  • —data/windows_nominal_power.jsonl: primary window dataset excluding the suspected power-limited tail.
  • —data/windows_suspected_power_limited.jsonl: late-run/power-constrained window dataset.
  • —raw/attempts.jsonl, raw/windows.jsonl, raw/events.jsonl: original run logs.
  • —raw/mactop_samples.csv.gz: compressed system-level DRAM/GPU/CPU/power samples.
  • —raw/server_logs.tar.gz: compressed llama-server logs.

Caveats

  • —mactop bandwidth is system-level telemetry, not per-process attribution.
  • —The suspected power-limited cohort should be used for sensitivity checks, not primary averages.
  • —This is a local-hardware inference benchmark, not a model quality benchmark.