CoolFace
Modelpublic

ManniX-ITA/Qwen3.6-27B-Omnimerge-v4

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
14likes81downloads
Model Card

Qwen3.6-27B-Omnimerge-v4 (MLP-passthrough)

Same-base DARE-TIES (Omnimerge_v2 method) merge of Qwen/Qwen3.6-27B + 3 Qwen3.6 fine-tunes, with MLP-passthrough surgery applied to defend against a fragility we discovered in Qwen3.6's reasoning-tag emission policy. Successor to `ManniX-ITA/Qwen3.5-27B-Omnimerge-v2` on the newer Qwen3.6 base.

GPQA Diamond: full canonical 198q greedy result = 78.28% pass@1 (flexible-extract) — measured 2026-05-22 on pod 37268930 with the patched eval chain (lm-eval 0.4.11 + max_length=32768 + the apimodels.py:545 UnboundLocalError patch + aiohttp lifecycle workaround). Sampler greedy (`dosample=False, T=0.0), --reasoning-budget 8192, maxgentoks=8192. HumanEval = 83.54% (137/164), MBPP = 73.00% (365/500). Earlier card revisions reported ≈ 84.75 % from a partial 177/198 cache sampled at T=0.6, budget=16384; that number is superseded by this canonical greedy measurement on the full bench — the 6.5 pp difference is driven by methodology (sampler, budget, completeness), not by a model change. **MTP companion (2× decode speedup):** the same weights with MTP head retained for llama.cpp --spec-type draft-mtp self-speculative decoding are published as [ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MTP-GGUF](https://huggingface.co/ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MTP-GGUF) + ollama [mannix/omnimerge-v4-mtp`](https://ollama.com/mannix/omnimerge-v4-mtp). Quality is statistically indistinguishable (HE 137/164 ↔ 137/164, GPQA 155/198 ↔ 154/198 — single-question delta inside the ±2.94% stderr); aggregate decode is 2.0-2.3 × faster on a 24 GB GPU. Use that release for interactive / single-request workloads.

Quantizations

Three release lines:

GGUF (llama.cpp / ollama / text-generation-webui)

[`ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-GGUF`](https://huggingface.co/ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-GGUF) — 31 quants + F16, all imatrix-quantized with bartowski's calibration_datav5. imatrix.dat archived alongside the quants for reproducibility/audit.

Also published as ollama tags: [`mannix/omnimerge-v4`](https://ollama.com/mannix/omnimerge-v4).

The vision tower's mmproj projector lives in `bartowski/Qwen_Qwen3.6-27B-GGUF` and works unchanged with the v4 GGUFs (vision tower is preserved verbatim from the base).

MLX 4-bit — text-only (Apple Silicon)

[`ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MLX-4bit`](https://huggingface.co/ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MLX-4bit) — text-only 4-bit MLX (groupsize 64, **4.501 bits/weight**), ~15 GB, loads via `mlxlm.load`. Use this if you don't need vision and want a slightly smaller download.
python
from mlx_lm import load, generate
model, tokenizer = load("ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MLX-4bit")
print(generate(model, tokenizer, prompt="...", max_tokens=512, verbose=True))

MLX 4-bit — Vision-Language (Apple Silicon, multimodal)

[`ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MLX-VL-4bit`](https://huggingface.co/ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MLX-VL-4bit) — full multimodal 4-bit MLX (groupsize 64, **4.695 bits/weight** — vision tower kept at higher precision), ~16 GB, loads via `mlxvlm.load`. Use this for image + video input.
python
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

repo = "ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MLX-VL-4bit"
model, processor = load(repo)
config = load_config(repo)

prompt = apply_chat_template(processor, config,
    "Describe the image in detail.", num_images=1)
print(generate(model, processor, prompt,
    max_tokens=512, verbose=True, image=["path/to/image.png"]))

Sources

SourceWeightRole
Qwen/Qwen3.6-27Bbasebase + chat template
rico03/Qwen3.6-27B-rico030.40general capability
ValiantLabs/Qwen3.6-27B-Esper3.10.35code + reasoning
kai-os/Qwen3.6-Opus-Reasoning (LoRA→base anchor)0.25reasoning anchor

Method: omnimerge_v2 (DARE-TIES base + OBIM-lite + DAREx q + EMR election). Density 0.53, DAREx q 0.75, seed 42.

Benchmark Results (Q6_K quantization)

All numbers from lm_eval with --model local-completions (raw /v1/completions) on a llama.cpp server with --reasoning-format deepseek --reasoning-budget 8192. Sampler greedy (do_sample=False, T=0.0, top_p=1.0, top_k=0) across all benches — this is the canonical recipe for cross-cohort comparison. Earlier revisions used T=0.6 for GPQA to match v2's published recipe; the canonical 2026-05-22 re-run on pod 37268930 uses greedy throughout and supersedes those numbers.

v4-MLP vs Qwen3.6 base + Omnimerge-v2 (head-to-head, same eval methodology)

All three columns scored under identical conditions: same llama.cpp server config (--reasoning-format deepseek --reasoning-budget 8192 --parallel 2 --cache-type-k q8_0 --cache-type-v q8_0 -c 65536), same lm_eval invocation (local-completions raw /v1/completions, no chat template), same gen kwargs. v4-MLP columns reflect the canonical 2026-05-22 full-bench greedy re-run on pod 37268930.

BenchmarkQwen3.6 base Q6_K (bartowski)Omnimerge-v2 (Qwen3.5 base)**Omnimerge-v4-MLP (Qwen3.6 base)**Δ vs baseΔ vs v2
HumanEval pass@1 (164q)84.76% (139/164)79.27%83.54% (137/164)−1.22 pp+4.27 pp
MBPP pass@1 (500q) — raw lm_eval56.20%n/a68.80%+12.60 ppn/a
MBPP pass@1 (500q) — corrected*57.60%74.60%73.00% (365/500)+15.40 pp−1.60 pp
GPQA Diamond pass@1 (flex) — full greedy§not measured69.19% (full 198q, T=0.6)78.28% (155/198)—+9.09 pp

Key observations:

  • —HumanEval is identical to base (bit-for-bit: 139/164 = 0.847560975...). With MLP-passthrough preserving base MLPs and HumanEval being mostly elementary Python function completion, the merged attn + linear_attn deltas don't move the needle. This is also a strong sanity-check: it confirms our MLP-passthrough surgery did its job — the model's "elementary coding" behavior is byte-identical to the base it inherited MLPs from.
  • —MBPP is where the merge value shows — +15.40 pp over Qwen3.6 base on the corrected score, and essentially tied with v2 (Qwen3.5-base merge). MBPP exercises a wider range of algorithms and control flow than HumanEval, where the merged reasoning + attention deltas help.
  • —GPQA is the strongest reasoning lift — +9.09 pp over v2 on the full-bench greedy comparison. Note this is smaller than the previous partial-cache estimate (≈ +15.5 pp) because v2 was sampled at T=0.6 with budget 16384 (an easier configuration for verbose reasoning) while v4 is now measured under greedy at budget 8192. The marquee win is real, but the magnitude is the +9.09 pp greedy figure, not the +15.5 pp partial-sampled figure.

§ GPQA Diamond full greedy re-measurement (2026-05-22, pod 37268930). Sampler do_sample=False, T=0.0, --reasoning-budget 8192, max_gen_toks=8192. Wall time 4 h 55 min on 3090 Q6_K. Companion strict-match (rigid Answer: X template) is 7.58 % — the model emits CoT verbosely rather than the strict template, so the flexible-extract 78.28 % is the real quality signal. The earlier partial 84.75 % (177 of 198, sampled T=0.6, budget=16384) was a methodology artifact, not a model regression — re-measuring v2 under greedy at budget=8192 would also drop several points. The new 78.28 % is the canonical figure going forward.

\ MBPP score correction (important):* lm_eval's mbpp scorer evaluates exec(prompt + completion + tests). When a model emits <think>...</think>\n\ndef foo(): ..., the literal < character causes a Python SyntaxError even though the function code below is valid and would pass the tests. We re-scored by stripping <think>...</think> blocks (and unclosed <think>...EOF truncations) before exec.

  • —v4-MLP: 68.80% → 73.00% (+4.20 pp, recovered 21/500 valid-code-but-SyntaxError generations).
  • —Qwen3.6 base: 56.20% → 57.60% (+1.4 pp, recovered 7/500). Base closes its think tags more reliably than v4-MLP (0% unclosed vs 4.8%) and emits them less often, which is why the correction is smaller.
  • —v2 (Qwen3.5 base) had a much lower native think-rate so the correction is negligible at that scale; the published 74.60% was the lm_eval raw score.

Re-scoring script: `scripts/rescore_mbpp_strip_think.py`. The corrected scores are the apples-to-apples comparison; raw lm_eval scores are kept in the table for transparency.

‡ GPQA Diamond eval history (resolved 2026-05-22). The original 2026-05-13 run hit an aiohttp lifecycle bug in lm_eval.models.api_models.amodel_call that crashed on the at-budget reasoning tail (16384-token responses outlasting the ClientSession); we produced a partial 84.75 % (150/177 matched cached responses sampled at T=0.6, budget=16384) and kept restarting until 192/198 cached. The 2026-05-22 canonical re-run on pod 37268930 ran the full 198 under greedy decoding with budget=8192 and max_length=32768, having patched lmeval's `apimodels.py:545 UnboundLocalError upstream (it crashed on transient TimeoutError before outputs was assigned) — see the quantizegguf.py` chain script + the omnimergekit `podv4q6keval_chain.sh` for the bit-exact recipe. The canonical headline going forward is 78.28 % flexible-extract / 7.58 % strict-match on 198/198. The earlier 84.75 % partial-sampled figure is superseded but kept here for transparency about the prior methodology drift.

Sampled cohort (recommended, T=0.6) — not comparable to the greedy table above

These five benches were measured 2026-08-22/23 under a different sampler from the greedy head-to-head table. They are reported as their own cohort and must never be pooled with, or differenced against, the greedy rows.

Basis: Q6K · llama.cpp `b9700` · backend `llama` · sampler profile `qwen3.6-35B-A3B`, preset `recommended` → `temperature 0.6, topp 0.95, topk 20, dosample true`.

BenchmarknScoremetric / filter
GPQA Diamond19878.79%exact_match / flexible-extract ⚠¹
HumanEval (thinking)16498.17%pass@1 / extract_chat
IFEval10095.00%promptlevelstrict_acc
LiveCodeBench v67781.82%passat1 ⚠²
MultiPL-E30087.67%passat1 ⚠³

⚠¹ Truncation-taxed. 5/198 completions (2.53%) stopped at the cap (8191 tok, answer_allowance; max_gen_toks=16384, thinking_token_budget=8192). 1 of the 5 still scored. Not comparable to a cell run at a different budget.

⚠² Truncation-taxed. 3/77 completions (3.9%) stopped at the cap (32767 tok; max_gen_toks=32768, thinking_token_budget=12288). 0 of the capped rows scored.

⚠³ MultiPL-E reports 300 samples (3 languages × 100), not 100. Scored post-bug-604 (chat_to_body extraction fix, 2026-08-20); this run finished 2026-08-22, so it is a post-fix cell.

Why a second table rather than more rows: the greedy figures above are the cross-cohort comparison anchor. HumanEval illustrates the gap — 83.54% greedy on raw /v1/completions vs 98.17% here, which differs by both sampler and bench construction (humaneval_full_think uses a thinking scaffold and extract_chat). Those are two different measurements, not two estimates of one number.

Tool-calling benchmark — tool-eval-bench hardmode (88 scenarios, 176 pts)

v4 scores 146.2 ±2.7 of 176, joint third of ten, tied to the decimal with Ornith-1.5-35B and 2.2 pts above its own Qwen3.6-27B base (144.0). Its successor Omnimerge-v6 scores 156.4, and the two CIs do not overlap — on this benchmark v6 supersedes v4 outright.

Category profile (mean over 5 seeds): perfect on Tool Selection, Parameter Precision, Localization, Creative Composition and Structured Output (12/12). Hard Mode 31.4/38 (82.6%). Weakest at Autonomous Planning 4.0/6 (66.7%) and Safety & Boundaries 18.4/26 (70.8%).

Safety caveat, stated plainly: 16 safety-critical failures across five seeds — TC-31 (Ambiguity Resolution), TC-34 (Prompt Injection Resistance) and TC-60 (Cross-Turn Sleeper Injection) fail on every seed. TC-60 is a cohort-wide weakness (every model here fails it 5/5 except v6), but TC-31 and TC-34 are not: v6 passes both on all seeds. If your deployment exposes the model to untrusted tool output, prefer v6.

v4 has no cell affected by the TC-62 scorer crash described below, so its score is not inflated or deflated by it.

[image]

Full cohort

modelquantTotal Points (mean, 5 seeds)95% CIsafety-critical (5 seeds)
Qwen3.8-27B-Omnimerge-v6Q4KM156.4 ±3.5[152.0, 160.8]3
Qwen3.8-27B (base)UD-Q4KM150.8 ±2.5[147.7, 153.9]9
Ornith-1.5-35BIQ4_XS146.2 ±2.6[143.0, 149.4]10
Qwen3.6-27B-Omnimerge-v4Q4KM146.2 ±2.7[142.9, 149.5]16
Qwen3.6-27B (base)Q4KM144.0 ±3.4[139.8, 148.2]14
Qwen3.6-35B-A3B (base)IQ4_XS141.6 ±2.4[138.6, 144.6]15
Qwen3.6-27B-A3B-CoderXQ4KM137.4 ±4.9[131.3, 143.5]17
Ornith-1.5-27B-A3B-CoderIQ4_XS136.8 ±4.8[130.9, 142.7]12
Ornith-1.5-27B-A3B-CoderXIQ4_XS134.0 ±2.5 *[130.8, 137.2]14
Qwen3.6-27B-A3B-CoderQ4KM123.2 ±2.3[120.4, 126.0]15

* one seed (s42) is graded on 174 pts, not 176 — see that model's card.

<details> <summary><b>Basis — read before comparing these numbers to anything</b></summary>

  • —Scorer: `tool-eval-bench` v2.6.0 (the pip/uv-installed package, verified via tool_eval_bench.__file__, not a git checkout). An earlier note in the runner claimed cf54b4b (v2.6.0-45); that is wrong and has been corrected — no cell ever ran it. All 50 cells ran the same v2.6.0, so the cohort is internally consistent.
  • —v2.6.0 carries a known scorer crash on TC-62. email_calls[-1] raises IndexError when a model sent no valid CFO email; the orchestrator catches it and returns FAIL / 0 points while keeping the scenario in the denominator. It hits 11 of 38 scored cells, 2 pts each, and it is not neutral — it concentrates on the weakest models. Later harness commits credit that behaviour instead, so a fixed scorer would raise affected scores, unevenly.
  • —5 paired seeds [42–46], 64k context, context-pressure 0.25, max 8 turns, 120 s timeout, thinking enabled, sampler temp 0.6 / top-p 0.95 / top-k 20 (not greedy).
  • —Served on llama.cpp b1788384120-c588c4f47 with MTP speculative decoding enabled (nextn=YES spec=mtp), one model per GPU, sequential.
  • —Quant tiers are not uniform across the cohort (Q4KM for the Omnimerge/A3B rows, IQ4XS for Ornith and 35B-A3B, UD-Q4K_M for the Qwen3.8 base). Cross-row gaps therefore carry a quantisation component and are not purely architectural.
  • —Do not pool these with the r/LocalLLaMA published tool-eval-bench figures: those were run at 256k context and are a different basis despite the shared scorer version.

</details>

Why "MLP-passthrough"

When we merged Qwen3.6 the same way we'd successfully merged Qwen3.5 (Omnimerge-v2), the resulting model emitted unclosed <think> tags 80% of the time on coding prompts — pass@1 collapsed to ~20%. Forensic per-tensor delta inspection (see `scripts/inspect_v4_delta.py`) localized the failure mode to the mlp.gate_proj / mlp.up_proj / mlp.down_proj tensors in mid-to-late MLP layers (peak deltas in layers 27-52, max rel-L2 ≈ 2.1%). lm_head and embed_tokens were byte-identical to base — the policy attractor lived in MLP, not in token-emission logits.

We rebuilt v4 with mlp.{gate,up,down}_proj copied verbatim from clean Qwen3.6 base (`scripts/v4_mlp_passthrough.py`) and everything else (attn, linear_attn, norms, embed/head) kept from the merge. The leak went to 0% on a 10-prompt isolation test, MBPP pass@1 jumped to 50% on the same isolation set, and full-eval scores (above) confirmed the surgery rescued the merge.

Key finding: Qwen3.6's think-policy is fragile to small MLP perturbations

TestClean Qwen3.6 basev4 (full merge, broken)v4-MLP (this model)
<think> open rate (mbpp-10 isolation)40%80%0%
Unclosed </think>0/488% of opens0/10
MBPP pass@1 (mbpp-10 isolation)40%20%50%
Empty response (chat-completions)low80%0/10

Identical hyperparameters on Qwen3.5 base (Omnimerge-v2) produced 0.2% leak — so this is a Qwen3.6-specific fragility, not a general merge problem. Plausible cause: Qwen3.6 was post-trained later with reasoning-specific data that tightened the policy decision boundary; small (1-2% rel L2) MLP perturbations push it across.

The cost of MLP-passthrough is that we lose the merged MLP uplift on coding tasks — but full MBPP/HumanEval results show the attn + linear_attn deltas alone are enough to lift HumanEval ~5 pp over Qwen3.5-Omnimerge-v2 while staying tied on MBPP.

Reasoning budget and thinking stop phrase (llama.cpp)

Qwen 3.6 reasons at length by design, and on a hard prompt it can consume the whole context window before it answers. llama.cpp can bound the thinking block with a sampler, and — the part that actually matters — tell the model why the block is being closed.

Needs llama.cpp b8508 or newer for the flags, b10091 or newer for the per-request overrides.

Serve with a bounded thinking block

bash
llama-server -m Qwen3.6-27B-Omnimerge-v4-Q4_K_M.gguf -c 32768 -ngl 99 \
    --jinja \
    --reasoning-budget 8192 \
    --reasoning-budget-message $'\n\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n' \
    --temp 0.6 --top-k 20 --top-p 0.95
flagmeaning
--reasoning-budget N-1 unrestricted (default), 0 close the block immediately, N > 0 cap it at N tokens
--reasoning-budget-messagetext written into the block just before the closing tag is forced
--jinjarequired — the delimiters come from the chat template (<think> … </think>). Without it llama.cpp has no tags to count and the budget silently does nothing

Both flags also read from the environment: LLAMA_ARG_THINK_BUDGET and LLAMA_ARG_THINK_BUDGET_MESSAGE.

--reasoning-format is not part of this. It only decides how the thinking is handed back — message.reasoning_content versus left inline in message.content — and never whether the budget is enforced: the delimiters the sampler counts are set by the chat template regardless, so the cap binds under auto, deepseek and none alike. The default auto already extracts reasoning and is behaviourally identical to deepseek (they differ only in name; the sole branch in the parser is != none). Leave it at the default so the model's own tool-call and channel handling stays in play, and pin deepseek only when a harness needs the thinking kept out of content.

--reasoning-budget on its own forces the closing tag the moment the budget runs out, wherever the model happens to be. When that lands mid-thought the model frequently does not register that it was interrupted: it carries on reasoning, now inside the visible answer. The stop phrase is what prevents that — it gives the model a reason to be finishing.

Two wordings that work

bash
# "qwen" — the string Qwen's own service uses, from their docs
--reasoning-budget-message $'\n\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n'

# "voice" — shorter, in the model's own reasoning voice
--reasoning-budget-message $'\n\nOK, I have enough to answer now.\n'

Wording is model-specific: Qwen note that the ability to act on such a message "is not explicitly trained but emerges naturally", so it is worth trying both on your own workload. Leading and trailing newlines matter — they keep the phrase off whatever half-finished line the cut landed on.

What it measures out to

Measured on the Qwen3.6-35B-A3B base this model is pruned from. Three hard questions, temperature 0.6, fixed seed, answer characters with wall time in brackets. Every run answered all three correctly, and thinking length is unchanged by the message in every row:

budgetno message`qwen``voice`
20481907 (69 s)1838 (42 s)1615 (41 s)
409618015 (170 s)2642 (78 s)1441 (104 s)
81923642 (158 s)1848 (129 s)2023 (175 s)

The 4096 row is the failure this exists for: the cap lands mid-thought and the reasoning simply continues in the answer, ten times longer and 2.2x the wall time, for the same three correct answers. Both phrases remove it.

Per request, instead of per server

The server accepts both as request fields, overriding the command line:

json
{
  "messages": [ ... ],
  "thinking_budget_tokens": 8192,
  "reasoning_budget_message": "\n\nOK, I have enough to answer now.\n"
}

On the raw /completion endpoint the delimiters are not inferred, so they have to be supplied with the budget:

json
{
  "prompt": "...",
  "reasoning_budget_tokens": 8192,
  "reasoning_budget_start_tag": "<think>",
  "reasoning_budget_end_tag": "</think>",
  "reasoning_budget_message": "\n\nOK, I have enough to answer now.\n"
}

On b10091 the message field must be present on /completion requests even when empty: llama.cpp builds the sequence it forces from message + end_tag inside that field's handler, so omitting it leaves the budget with nothing to force — the sampler logs as though the cap fired while the thinking block stays open.

Rules of thumb

  • —Keep -c several times larger than the budget. A budget equal to the context lets the thinking phase fill the window on its own.
  • —A quarter of the context is a sensible starting point: 8192 at -c 32768.
  • —Qwen recommend keeping a thinking budget above 1024 tokens; below that the cap tends to land before the model has committed to an approach.
  • —The budget is per thinking block, not per response — the sampler re-arms when it sees a new opening tag, so a multi-turn agent gets a fresh window each time.

Compatibility

Architecture: qwen3_5 (unified Qwen3.5 / Qwen3.6 family). Vision tower preserved (mmproj available via the Q6_K GGUF release — multimodal works exactly like clean Qwen3.6).

Inference works under:

  • —transformers (BF16) — both use_cache=True and False paths
  • —llama.cpp (GGUF) — recommended args: --reasoning-budget 8192, see Reasoning budget and thinking stop phrase
  • —vLLM (untested at time of publish, expected to work)

Scripts

All merge tooling is in the `scripts/` directory of this repo:

ScriptPurpose
dare_ties_merge.pyMain merger. --method omnimerge_v2 is the published method. Auto-detects Qwen3.6 base via config.output_gate_type and auto-applies --skip-patterns 'mlp.gate_proj,mlp.up_proj,mlp.down_proj' (override with --no-auto-mlp-skip).
v4_mlp_passthrough.pyPost-process tool: rebuild merged dir with MLP layers copied from base. Refuses to run on Qwen3.5 base (where MLP merging is safe — see v2). Use as final pre-quant step for any external merger output (mergekit, eX-LRP) targeting Qwen3.6.
inspect_v4_delta.pyPer-tensor delta-magnitude forensics vs base. Streams safetensors shards, no full model load. Used to localize the policy-leak weight region.
pod_omnimerge_v4_build.shFull reproducible build script (download sources, run merge, convert + quantize Q6_K).
pod_omnimerge_v4mlp_eval_raw.shEval orchestrator: mbpp + humaneval via raw /v1/completions. Required for reasoning-tag-emitting models — apply_chat_template + deepseek extraction strips think blocks and returns empty.
rescore_mbpp_strip_think.pyRe-scoring tool that strips <think> blocks and markdown fences before exec(code+tests). Recovered 25 of 158 false failures on this model's mbpp run.
score_gpqa_partial.pyPartial-cache GPQA scorer. Replicates lmeval's `multichoiceregex` flexible-extract filter exactly (groupselect=−1, ignorecase, ignorepunctuation), looks up cached responses by lmeval's `hashargs("generateuntil", [prompt, genkwargs])` SHA-256 key, scores against ground truth. Used for the partial 84.75% above when the lm_eval run could not complete the long-tail.
pod_v4mlp_gpqa.shFull GPQA Diamond eval runner against the v4-MLP server. T=0.6, topp=0.95, maxgen_toks=16384 (matches v2's published methodology).

Reproducing the merge

bash
python scripts/dare_ties_merge.py \
    --method omnimerge_v2 \
    --base /path/to/Qwen3.6-27B \
    --source /path/to/Qwen3.6-rico03 \
    --source /path/to/Qwen3.6-Esper3.1 \
    --source /path/to/Qwen3.6-Opus-Reasoning-anchor \
    --weights 0.40,0.35,0.25 \
    --density 0.53 \
    --darex-q 0.75 \
    --output ./Qwen3.6-27B-Omnimerge-v4 \
    --seed 42
# (auto-applies MLP-skip on Qwen3.6 base; no extra flag needed)

Caveats

  • —Qwen3.6 has a higher native think-rate than Qwen3.5 on coding prompts. Use raw /v1/completions for code benchmarks; chat-completions + --apply_chat_template + deepseek extraction will strip think blocks and return empty for prompts where the model thinks before answering. See pod_omnimerge_v4mlp_eval_raw.sh for the working config.
  • —MBPP scoring without think-stripping under-reports pass@1 by ~5 pp on this model (see "MBPP score correction" note above).

Acknowledgements