GuYueshou/SenseNova-U1.5-8B-MoT-Q4_K_S-L0DownFix-GGUF
SenseNova U1.5 Final Q4KS - Layer-0 Down Projection Fix
Community repair / derivative GGUF. Not an official SenseNova or RealRebel release.
Community repair of RealRebel's Q4KS artifact for the Final release of sensenova/SenseNova-U1.5-8B-MoT.
TL;DR
- The original RealRebel Q4KS artifact produced severely broken output under a verified 50-step, CFG 4, no-LoRA Final baseline.
- Causal intervention localized the main failure to one task-sensitive tensor:
language_model.model.layers.0.mlp.down_proj.weightin the layer-0 understanding branch. - This repaired GGUF restores only that tensor to the exact official Final BF16 values, losslessly widened into an F32 GGUF tensor container for compatibility with the existing loader.
- Normal semantic and structural image generation was restored under the previously failing 50-step baseline.
- All other 1,115 tensor payloads are unchanged; tensor count, tensor names, tensor order, and GGUF metadata are preserved.
What was wrong?
The source file was not generally corrupted, and this investigation did not show that the Q4K dequantizer is broken. Tensor names and shapes were intact, the Q4K decode path agreed with an independent implementation in the audited samples, and the failure reproduced without LoRA.
The issue was localized to the layer-0 understanding mlp.down_proj. Its Q4K values are not obviously poor in global weight space. In fact, against the official Final BF16 tensor, the source Q4KS tensor had lower global relative L2 error than the working Q40 behavioral control:
However, global weight-space accuracy did not predict behavior on the model's measured activation subspace.
Root cause
The observed causal chain was:
layer-0 understanding mlp.down_proj quantization
-> prefix hidden-state drift
-> prefix V-cache corruption
-> generation attention receives incorrect cached values
-> generation trajectory fails
-> broken image outputAt layer 5, the unmodified Q4KS prefix path had already diverged severely from the working Q4_0 behavioral control:
This single-tensor intervention therefore did more than postpone divergence by one block: it prevented the early-prefix failure from exploding through layers 1-5.
Final-grounded layer-0 intervention
Using the official Final layer-0 projections as a local reference, the normalized layer-1 state showed the following relative errors:
Replacing only the Final down projection reduced the baseline error by about 61.22%, while replacing the attention projections alone had little effect. This identified the layer-0 understanding mlp.down_proj as the first major task-sensitive quantization bottleneck in the measured path.
Why Q40 could work while Q4K_S failed
This result should not be interpreted as "Q40 is a better quantization format" or "Q4K is generally worse."
The Q4KS tensor was closer to the official Final tensor under global weight-space metrics. Yet, on two measured real activation inputs, the Q40 operator happened to produce lower output error than Q4K_S:
The relevant failure mechanism is therefore quantization-error geometry relative to the real activation subspace, not merely the average magnitude of weight error.
Global weight-space accuracy does not necessarily predict activation-weighted operator accuracy.
The fix
Exactly one tensor was changed:
language_model.model.layers.0.mlp.down_proj.weight
Q4_K -> F32 container holding exact official Final BF16 valuesMore precisely:
Official Final BF16 values are losslessly widened into an F32 GGUF tensor container for compatibility with the existing loader.
The F32 container is a storage and loader-compatibility decision, not a claim that the original model was trained with FP32 weights or that the whole model was converted to FP32.
The existing REBEL loader materializes GGUF F16/F32 tensors as regular floating Linear modules, while Q4_K tensors use GGUFDequantLinear. Widening BF16 to F32 is exact; converting this tensor to F16 changed 2,638 BF16 values in the tested round trip. Using F32 therefore avoids extra rounding and requires no loader modification.
Artifact integrity
The repaired GGUF is 13,228,580,224 bytes. Its net growth over the source Q4KS is 173,015,040 bytes (165 MiB).
Repaired GGUF SHA-256
393ccdb34f1173914a818c4b8040bc0978a618004da801fabd2474c707c990aeThe included build report and payload manifest contain the detailed qtype inventory, loader checks, per-tensor payload hashes, and aggregate manifest hash.
About the loader's missing/unexpected key counts
The source loader baseline reports 588 missing and 588 unexpected post-materialization wrapper-state keys. The repaired artifact reports 587 / 587 because the repaired tensor is now a regular floating Linear rather than a GGUF wrapper. This is one fewer wrapper-key pair, not a new model-loading failure. No new missing keys, unexpected keys, or meta tensors were introduced.
Validation
Original-artifact failure baseline
The original Q4KS failed under this baseline:
Final metadata
no LoRA
50 steps
CFG 4
cfg_norm none
timestep shift 3
seed 123456789
2048 x 2048
SDPA
prefetch count 1
existing 11 official-Final BF16 overridesThe original artifact produced a severely abnormal white/granular output with broken image structure.
Permanent-GGUF prepublication acceptance
The permanent artifact SenseNova-U1.5-8B-MoT-Q4_K_S-L0DownFix.gguf was then loaded and run directly under the same baseline. This acceptance run used:
- the permanent repaired GGUF itself;
- no additional layer-0 runtime intervention (
runtime_intervention: null); - no LoRA;
- the same Final metadata, prompt, seed, 2048x2048 resolution, 50 steps, CFG 4,
cfg_norm=none, timestep shift 3, SDPA backend, prefetch count, and existing 11 official-Final BF16 overrides as the failing baseline.
The repaired target tensor materialized as a regular Linear, and its loaded value converted to BF16 matched the official Final BF16 tensor exactly. The run completed successfully and restored normal semantic and structural image generation: the person, dragon, butterfly, stream, mountains, sun, and their main spatial relationships were recovered.
As an artifact-equivalence check, the permanent-GGUF output was pixel-exact against the earlier single-tensor runtime-intervention validation output:
The detailed prepublication acceptance record is included as SenseNova_U1.5_Q4KS_L0DownFix_50step_Base_00001__report.json.
This is intentionally described as restoration of normal semantic and structural generation. It is not a claim of visual or bit-exact equivalence to the complete official BF16 model.
Requirements and tested integration
This GGUF is not a standalone replacement for the official model repository. The validated setup used:
- official Final metadata from `sensenova/SenseNova-U1.5-8B-MoT`;
- the existing `ComfyUI_SenseNova_U1_REBEL` integration;
- the existing 11-tensor official-Final BF16 override behavior used by that setup;
- the repaired GGUF loaded through the REBEL SenseNova model loader.
The local validation launch configuration used environment variables equivalent to:
set "SENSENOVA_U1_REPO=X:\path\to\official-final-metadata"
set "SENSENOVA_GGUF_BF16_OVERRIDE=X:\path\to\official-final-shards\model-00001-of-00008.safetensors"
set "SENSENOVA_U1_DISABLE_PINNED_OFFLOAD=1"
python ComfyUI\main.py --windows-standalone-buildReplace the example paths with your own local paths. No machine-specific launch script, proxy setting, token, cookie, or credential is included in this repository.
Limitations
- Verified only on the tested SenseNova U1.5 Final setup and integration described above.
- This repairs one specific RealRebel Q4KS artifact; it does not prove that Q4K is generally inferior to Q40.
- It does not prove that every SenseNova Q4_K conversion has the same failure.
- Only selected validation prompts, seeds, and settings were tested.
- It is not an official SenseNova or RealRebel release and is not endorsed by either project.
- It is not claimed to be bit-exact or quality-equivalent to the complete official BF16 model.
- Existing upstream integration and external official metadata/BF16 override requirements still apply.
Provenance
- Official Final model: `sensenova/SenseNova-U1.5-8B-MoT`
- Source GGUF repository: `realrebelai/SenseNova-U1.5-8B_GGUFs`
- Source artifact: `SenseNova-U1.5-8B-MoT-Q4_K_S.gguf`
- Repair repository:
GuYueshou/SenseNova-U1.5-8B-MoT-Q4_K_S-L0DownFix-GGUF
The source Q4KS artifact and the official Final shard used for repair were kept unmodified during construction and validation.
License
The official SenseNova Final repository and the RealRebel GGUF repository both currently declare the Apache License 2.0. This derivative is released under Apache-2.0, subject to the upstream license terms and notices.
