CoolFace
Modelpublic

GuYueshou/SenseNova-U1.5-8B-MoT-Q4_K_S-L0DownFix-GGUF

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes300downloads
Model Card

SenseNova U1.5 Final Q4KS - Layer-0 Down Projection Fix

Community repair / derivative GGUF. Not an official SenseNova or RealRebel release.

Community repair of RealRebel's Q4KS artifact for the Final release of sensenova/SenseNova-U1.5-8B-MoT.

TL;DR

  • —The original RealRebel Q4KS artifact produced severely broken output under a verified 50-step, CFG 4, no-LoRA Final baseline.
  • —Causal intervention localized the main failure to one task-sensitive tensor: language_model.model.layers.0.mlp.down_proj.weight in the layer-0 understanding branch.
  • —This repaired GGUF restores only that tensor to the exact official Final BF16 values, losslessly widened into an F32 GGUF tensor container for compatibility with the existing loader.
  • —Normal semantic and structural image generation was restored under the previously failing 50-step baseline.
  • —All other 1,115 tensor payloads are unchanged; tensor count, tensor names, tensor order, and GGUF metadata are preserved.

What was wrong?

The source file was not generally corrupted, and this investigation did not show that the Q4K dequantizer is broken. Tensor names and shapes were intact, the Q4K decode path agreed with an independent implementation in the audited samples, and the failure reproduced without LoRA.

The issue was localized to the layer-0 understanding mlp.down_proj. Its Q4K values are not obviously poor in global weight space. In fact, against the official Final BF16 tensor, the source Q4KS tensor had lower global relative L2 error than the working Q40 behavioral control:

Tensor representationWeight-space relative L2 vs. official Final
RealRebel Q4KSabout 7.42%
Working Q4_0 controlabout 9.04%

However, global weight-space accuracy did not predict behavior on the model's measured activation subspace.

Root cause

The observed causal chain was:

text
layer-0 understanding mlp.down_proj quantization
-> prefix hidden-state drift
-> prefix V-cache corruption
-> generation attention receives incorrect cached values
-> generation trajectory fails
-> broken image output

At layer 5, the unmodified Q4KS prefix path had already diverged severely from the working Q4_0 behavioral control:

Layer-5 signalCosineRelative L2
Prefix hidden state, original Q4KSabout 0.0577about 656.80%
Prefix V cache, original Q4KSabout 0.1085about 813.22%
Prefix hidden state, Final layer-0 down projectionabout 0.9775about 21.42%
Prefix V cache, Final layer-0 down projectionabout 0.9774about 21.25%

This single-tensor intervention therefore did more than postpone divergence by one block: it prevented the early-prefix failure from exploding through layers 1-5.

Final-grounded layer-0 intervention

Using the official Final layer-0 projections as a local reference, the normalized layer-1 state showed the following relative errors:

Layer-0 pathLayer-1 normalized relative error
Q4KS baselineabout 8.1757%
Final attention projections onlyabout 7.6351%
Final mlp.down_proj onlyabout 3.1703%
Final MLP projectionsabout 2.5389%

Replacing only the Final down projection reduced the baseline error by about 61.22%, while replacing the attention projections alone had little effect. This identified the layer-0 understanding mlp.down_proj as the first major task-sensitive quantization bottleneck in the measured path.

Why Q40 could work while Q4K_S failed

This result should not be interpreted as "Q40 is a better quantization format" or "Q4K is generally worse."

The Q4KS tensor was closer to the official Final tensor under global weight-space metrics. Yet, on two measured real activation inputs, the Q40 operator happened to produce lower output error than Q4K_S:

Shared activation inputQ4_K_S down vs. FinalQ4_0 down vs. Final
aKabout 2.4074%about 1.8224%
a0about 2.4275%about 1.8125%

The relevant failure mechanism is therefore quantization-error geometry relative to the real activation subspace, not merely the average magnitude of weight error.

Global weight-space accuracy does not necessarily predict activation-weighted operator accuracy.

The fix

Exactly one tensor was changed:

text
language_model.model.layers.0.mlp.down_proj.weight
Q4_K -> F32 container holding exact official Final BF16 values

More precisely:

Official Final BF16 values are losslessly widened into an F32 GGUF tensor container for compatibility with the existing loader.

The F32 container is a storage and loader-compatibility decision, not a claim that the original model was trained with FP32 weights or that the whole model was converted to FP32.

The existing REBEL loader materializes GGUF F16/F32 tensors as regular floating Linear modules, while Q4_K tensors use GGUFDequantLinear. Widening BF16 to F32 is exact; converting this tensor to F16 changed 2,638 BF16 values in the tested round trip. Using F32 therefore avoids extra rounding and requires no loader modification.

Artifact integrity

CheckResult
Tensor count before / after1,116 / 1,116
Tensor namesExact match
Tensor orderExact match
GGUF metadataUnchanged
general.architecturePreserved as wan
Changed tensor payloadsExactly 1
Unchanged tensor payloads1,115 individually SHA-256 verified; 0 mismatches
Target qtypeQ4_K -> F32
Target runtime moduleRegular Linear, not GGUFDequantLinear
Target after loadBF16 SHA-256 exactly matches official Final BF16
Meta parameters / buffers0 / 0

The repaired GGUF is 13,228,580,224 bytes. Its net growth over the source Q4KS is 173,015,040 bytes (165 MiB).

Repaired GGUF SHA-256

text
393ccdb34f1173914a818c4b8040bc0978a618004da801fabd2474c707c990ae

The included build report and payload manifest contain the detailed qtype inventory, loader checks, per-tensor payload hashes, and aggregate manifest hash.

About the loader's missing/unexpected key counts

The source loader baseline reports 588 missing and 588 unexpected post-materialization wrapper-state keys. The repaired artifact reports 587 / 587 because the repaired tensor is now a regular floating Linear rather than a GGUF wrapper. This is one fewer wrapper-key pair, not a new model-loading failure. No new missing keys, unexpected keys, or meta tensors were introduced.

Validation

Original-artifact failure baseline

The original Q4KS failed under this baseline:

text
Final metadata
no LoRA
50 steps
CFG 4
cfg_norm none
timestep shift 3
seed 123456789
2048 x 2048
SDPA
prefetch count 1
existing 11 official-Final BF16 overrides

The original artifact produced a severely abnormal white/granular output with broken image structure.

Permanent-GGUF prepublication acceptance

The permanent artifact SenseNova-U1.5-8B-MoT-Q4_K_S-L0DownFix.gguf was then loaded and run directly under the same baseline. This acceptance run used:

  • —the permanent repaired GGUF itself;
  • —no additional layer-0 runtime intervention (runtime_intervention: null);
  • —no LoRA;
  • —the same Final metadata, prompt, seed, 2048x2048 resolution, 50 steps, CFG 4, cfg_norm=none, timestep shift 3, SDPA backend, prefetch count, and existing 11 official-Final BF16 overrides as the failing baseline.

The repaired target tensor materialized as a regular Linear, and its loaded value converted to BF16 matched the official Final BF16 tensor exactly. The run completed successfully and restored normal semantic and structural image generation: the person, dragon, butterfly, stream, mountains, sun, and their main spatial relationships were recovered.

As an artifact-equivalence check, the permanent-GGUF output was pixel-exact against the earlier single-tensor runtime-intervention validation output:

CheckResult
Permanent-GGUF PNG file SHA-25654915280ff9fea76bbd7397197d154d73a8e7e9b9eca802fef732a26b2010d91
Runtime-intervention PNG file SHA-256fc3c91d616690f0f86f2d734e7d99d86b1fd9bdf0af62eb3c8eabacccaf2e08f
Permanent-GGUF decoded pixel-buffer SHA-25654353adb2812506ae777afb1786725b842a37642b0ac12e1bf831917e9b527cd
Runtime-intervention decoded pixel-buffer SHA-25654353adb2812506ae777afb1786725b842a37642b0ac12e1bf831917e9b527cd
Pixels exactly equalYes
Maximum absolute pixel difference0

The detailed prepublication acceptance record is included as SenseNova_U1.5_Q4KS_L0DownFix_50step_Base_00001__report.json.

This is intentionally described as restoration of normal semantic and structural generation. It is not a claim of visual or bit-exact equivalence to the complete official BF16 model.

Requirements and tested integration

This GGUF is not a standalone replacement for the official model repository. The validated setup used:

The local validation launch configuration used environment variables equivalent to:

bat
set "SENSENOVA_U1_REPO=X:\path\to\official-final-metadata"
set "SENSENOVA_GGUF_BF16_OVERRIDE=X:\path\to\official-final-shards\model-00001-of-00008.safetensors"
set "SENSENOVA_U1_DISABLE_PINNED_OFFLOAD=1"
python ComfyUI\main.py --windows-standalone-build

Replace the example paths with your own local paths. No machine-specific launch script, proxy setting, token, cookie, or credential is included in this repository.

Limitations

  • —Verified only on the tested SenseNova U1.5 Final setup and integration described above.
  • —This repairs one specific RealRebel Q4KS artifact; it does not prove that Q4K is generally inferior to Q40.
  • —It does not prove that every SenseNova Q4_K conversion has the same failure.
  • —Only selected validation prompts, seeds, and settings were tested.
  • —It is not an official SenseNova or RealRebel release and is not endorsed by either project.
  • —It is not claimed to be bit-exact or quality-equivalent to the complete official BF16 model.
  • —Existing upstream integration and external official metadata/BF16 override requirements still apply.

Provenance

The source Q4KS artifact and the official Final shard used for repair were kept unmodified during construction and validation.

License

The official SenseNova Final repository and the RealRebel GGUF repository both currently declare the Apache License 2.0. This derivative is released under Apache-2.0, subject to the upstream license terms and notices.