CoolFace
Modelpublic

redashes/Qwen3.8-27B-BF16-SSMFIX

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
10likes936downloads
Model Card
πŸ“– δΈ­ζ–‡η‰ˆθ―΄ζ˜Ž β€” δΈ­ζ–‡ζ¨‘εž‹ε‘
⚠️ Experimental release β€” read Section 0 and the Disclaimer before use.

Qwen3.8-27B-BF16-SSMFIX (v2 Β· luffy per-layer Ξ±)

A conv1d-repaired Qwen3.8-27B: fixes the SSM scale-drift that silently degrades long-context generation.

Publisher's statement: I release this model not as a recommendation for use in daily life or work, but as a practical verification of a community hypothesis, and as a foundation platform for those interested in researching this field. All test data reflects verification within my personal capability; having more people validate it in more real-world scenarios will allow the truth of this theory to be tested faster and more authentically.

This model applies per-layer Ξ±-scaling to the anomalous linear_attn.conv1d.weight tensors in Qwen3.8-27B, following the methodology first disclosed by LuffyTheFox (Sig-ScaleSync) and independently re-implemented by FGDumitru (qwen-ssm-repair) β€” this release is the quantitative, community cross-validated proof that the fix works.

0. About This Release β€” an Independent Verification of the Community "Sig-ScaleSync" Investigation

This repository does not claim to be an official or definitive fix. It is a verification experiment around the community investigation first published by LuffyTheFox (Hugging Face: LuffyTheFox), who named his method Sig-ScaleSync (later folded into his broader "Genesis" pipeline). We replicated his core hypothesis independently β€” measuring conv1d weight-scale drift on the official Qwen3.8-27B weights, applying minimal per-layer Ξ± rescaling, and (unlike the original author) subjecting the repaired weights to a full controlled benchmark battery against the official baseline.

LuffyTheFox's original materials:

  • β€”Main model card, Genesis project (Qwen3.6-35B-A3B series): https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V8-GGUF
  • β€”*Direct analysis of this exact model, Qwen3.8-27B β€” discussion #38 "Why Qwen3.8-27B overthinks? Here the reason.":* https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V8-GGUF/discussions/38

His core thesis, in his own words ("Genesis" concept):

"During training, ALL models don't just learn knowledge – they also accumulate random noise in their tensors. This noise builds up and creates something I call the Noise Gate β€” a fundamental barrier that stops LLM models from learning further and makes them unstable, verbose, and prone to hallucinations." "LLM models often have: … Scale mismatches: one layer's weights are 10Γ— larger than its peers for no good reason …" "On first stage I scan ssm_conv1d tensors in model, they handle long context memory. I repair balance between heads in them." "My approach fixes all of that without retraining β€” pure numerical surgery on the raw bytes of the file."

He concluded with a strong claim about this exact model:

"That is also why I will not make Genesis for 27B. You cannot fix this by patching a few tensors or doing SVD to fix noise gate. The SSM input pathway is damaged across too many layers."

What this experiment adds

  • β€”His diagnosis confirms our independent measurement. The 8 layers we flagged (52/53/56/57/58/60/61/62) are identical to his Ξ±-based list, and our applied scale factors (0.481–0.653) match his Ξ± range (0.48–0.65).
  • β€”We tested, rather than asserted. We ran a full controlled battery (GSM8K, CMMLU, TruthfulQA, IFEval, MT-Bench) against the official baseline on identical hardware/stack. Results are in the Evaluation section below.
  • β€”Verdict vs. his "cannot fix" claim: partial refutation. A small tensor patch did move generative metrics substantially (TruthfulQA-gen +6~8pp) β€” but it also hurt closed-book knowledge (CMMLU βˆ’1.8pp) and slightly reduced conversational quality under the official MT-Bench protocol (βˆ’0.19 vs official, see Evaluation). So a few-tensor patch is not a free lunch: it trades a little knowledge and a little dialogue finesse for noticeably better generation/hallucination behavior.

This release is the measurable record of that experiment, not a recommendation to prefer it over the official weights. Use accordingly.

Why this model exists

Qwen 3.5/3.8 hybrid models mix full-attention layers with GatedDeltaNet SSM layers. The SSM recurrence is governed by 1D convolutional weights (linear_attn.conv1d.weight). In the official Qwen3.8-27B weights, 8 of the last layers have a significantly inflated conv1d std (vs. the ~0.042 sibling median):

LayerΞ± appliedpost-fix std
520.59010.0471
530.55480.0437
560.54490.0425
570.53570.0410
580.60970.0432
600.48140.0398
610.65330.0420
620.61860.0452

These layers are the same 8 flagged by LuffyTheFox (Ξ± 0.48–0.65) and overlap FGDumitru's detection (Ξ± 0.61–0.70) β€” independent implementations, convergent diagnosis. Without repair, the drifted scales let the recurrent state saturate/collapse: long-context (75k+) collapse, repetition loops, mid-generation truncation, and "philosophizing" drift where the model abandons the task. Short-context perplexity looks normal β†’ silent degradation.

Community Cross-Validation

SourceMethodAnomalous layersΞ± range
LuffyTheFox (HF discussions #38, Sig-ScaleSync/Genesis)Noise-gate theory, per-layer Ξ±Same 8 (52/53/56/57/58/60/61/62)0.48–0.65
FGDumitru (qwen-ssm-repair, MIT)MAD Z-score + peer-group median scalingOverlapping tail layers0.61–0.70
This release (v2)Per-layer strict Ξ± (Luffy method)Same 80.481–0.653

This release adopts the strict per-layer Ξ± from LuffyTheFox (not FGDumitru's median-normalization), because our full evaluation shows it preserves instruction-following and knowledge better (see table below). All weights are bit-exact except the 8 repaired tensors; model_type=qwen3_5 VLM integrity confirmed (visual / linear_attn / mtp intact).

Evaluation (vLLM, identical harness)

Metricofficial BF16v2 (per-layer Ξ±, this release)
MT-Bench avg8.798.60
IFEval prompt strict0.51940.5194
IFEval inst strict0.62470.6343
GSM8K strict0.96060.9644
CMMLU0.71790.6996
TruthfulQA mc1 / mc20.3647 / 0.54180.3758 / 0.5513
TQA gen rouge1/2/L, bleu0.284/0.162/0.280/0.1780.345/0.246/0.345/0.256

Takeaways:

  • β€”7 of 10 metrics β‰₯ or β‰ˆ official; the notable gaps are CMMLU (βˆ’1.8pp, knowledge-heavy) and MT-Bench (βˆ’0.19, conversational).
  • β€”TruthfulQA generation up +6~8pp across the board β†’ strong hallucination reduction (the main measurable win of the repair).
  • β€”MT-Bench (official protocol, per-category avg): v2 loses most on reasoning (βˆ’0.75), writing (βˆ’0.45), math (βˆ’0.30); gains on humanities (+0.30) and extraction (+0.15) β€” see the detailed MT-Bench section below.
  • β€”v1 (median norm) is deprecated and removed from this repo; v2 is the only SSMFIX variant shipped here.

MT-Bench β€” updated protocol results (2026-08-19)

⚠️ Supersedes the numbers published earlier. The previous MT-Bench scores on this card (7.05 / 7.15 / 7.47) came from a run with a broken harness: 1. max_model_len=8192 β†’ long reasoning-model answers retried at max_tokens=8192 overflowed and returned HTTP 400 β†’ the judge assigned fake 1.0 scores to ~1/5 of turns. 2. A single generic judge prompt was used for all categories, whereas the official FastChat protocol uses a dual-track judge: math/reasoning/coding are graded against the official GPT-4 reference answers (single-math-v1), all other categories use single-v1. 3. Thinking mode was ON (Qwen3.8 defaults to it). The official MT-Bench protocol assumes non-thinking chat models, so the old numbers were not comparable to official leaderboards. All three issues are fixed in this rerun: thinking OFF (enable_thinking=false), official per-category temperatures (math/coding/reasoning/extraction 0.0, stem/humanities 0.1, writing/roleplay 0.7), dual-track judge with GPT-4 references, and official turn1/turn2 aggregation. Judge: deepseek-v4-flash (temperature 0), 160/160 valid, zero failed turns. Use the numbers below; the old ones are void.
Categoryofficial BF16 (t1/t2/avg)v2 SSMFIX (t1/t2/avg)Ξ” (v2 βˆ’ official)
Overall8.96 / 8.61 / 8.798.94 / 8.26 / 8.60βˆ’0.19
writing9.10 / 8.30 / 8.708.90 / 7.60 / 8.25βˆ’0.45
roleplay9.00 / 8.60 / 8.808.40 / 8.70 / 8.55βˆ’0.25
reasoning9.80 / 9.20 / 9.509.50 / 8.00 / 8.75βˆ’0.75
math10.00 / 10.00 / 10.0010.00 / 9.40 / 9.70βˆ’0.30
coding9.00 / 8.20 / 8.609.10 / 8.00 / 8.55βˆ’0.05
extraction7.90 / 8.70 / 8.308.40 / 8.50 / 8.45+0.15
stem8.20 / 7.10 / 7.658.20 / 6.80 / 7.50βˆ’0.15
humanities8.70 / 8.80 / 8.759.00 / 9.10 / 9.05+0.30

Other metrics β€” why they remain valid (2026-08-19)

The five non-MT-Bench metrics (GSM8K, CMMLU, TruthfulQA, IFEval) run on a different chain than MT-Bench and were not hit by the three contamination mechanisms that voided the old MT-Bench numbers. The evidence below is verified against the actual run artifacts and code, not asserted.
  1. 1.Endpoint: raw `/v1/completions`, not chat. All five metrics go through lmeval's `local-completions` backend, which sends bare text-completion requests. MT-Bench alone uses `/v1/chat/completions` (chat template applied β†’ Qwen3.8's thinking mode ON by default), which was one of the old-MT-Bench contamination sources. The eval chain never applies the chat template; run logs show `huggingface tokenizer backend`, no `applychat_template` call.
  1. 1.Thinking is never triggered (tokenizer-verified). Qwen3.8 enters thinking mode only when the chat template injects the "Reasoning effort is set to xhigh…" system directive plus dedicated thinking tokens (248068/248069). We tokenized the real eval prompts with the actual Qwen3.8-27B tokenizer: bare completion prompts contain zero thinking tokens; only chat-template rendering does. So the eval runs are structurally thinking OFF β€” the same state as the corrected MT-Bench rerun, by construction.
  1. 1.Deterministic generation. lmeval's completions payload defaults to `temperature=0` (verified in `LocalCompletionsAPI.create_payload`). No sampling variance.
  1. 1.`max_gen_toks=2048` is uniform across every result referenced on this card. All result-file model_args snapshots (official / v2 for GSM8K / CMMLU / TQA / IFEval) carry max_gen_toks: 2048. The only run that ever used the 256-token default (an early port-8134 batch, source of the old "gsm8k fake drop") was superseded by -mg2048 reruns and is not referenced here.
  1. 1.Task-type immunity. CMMLU and TruthfulQA-mc1/mc2 are loglikelihood tasks (they score prompt probabilities, they do not generate); GSM8K / TQA-gen / IFEval are generative but run on the raw-completion chain above where thinking is structurally off. All three models were evaluated on identical chains, so every comparison on this card is apples-to-apples.

Conclusion: GSM8K / CMMLU / TruthfulQA / IFEval numbers on this card are trustworthy and protocol-consistent with the corrected MT-Bench rerun (thinking OFF, temperature 0, 2048-token budget).

Usage

python
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained("redashes/Qwen3.8-27B-BF16-SSMFIX", trust_remote_code=True)
processor = AutoProcessor.from_pretrained("redashes/Qwen3.8-27B-BF16-SSMFIX", trust_remote_code=True)

Provenance

  • β€”Base: official Qwen/Qwen3.8-27B BF16 (untouched except repaired tensors)
  • β€”Repair script: per-layer Ξ± on model.language_model.layers.<N>.linear_attn.conv1d.weight; atomic shard rewrites with .orig backups; 1199 keys verified, 48 conv1d keys verified, 0 remaining anomalous layers (ratio > 1.6)
  • β€”Method credit: LuffyTheFox (Sig-ScaleSync) / FGDumitru (qwen-ssm-repair)
  • β€”Produced by: hermes-nova

Disclaimer

  • β€”Weights are derived from the official Apache-2.0 release; the Apache 2.0 license is inherited.
  • β€”Only 8 conv1d tensors were rescaled; all other tensors are bit-identical to the official release.
  • β€”This model is an independent verification experiment of a community hypothesis (LuffyTheFox's Sig-ScaleSync, cross-validated by FGDumitru). Do not treat it as a production recommendation. Prefer the official weights unless you specifically need the generative-quality profile measured here.
  • β€”The original author's materials are linked in Section 0; any claims about his method are his own words, quoted verbatim.

License

Apache-2.0 (model weights follow the original Qwen license terms).