CoolFace
Modelpublic

chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-MLX-5bit

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes741downloads
Model Card

Qwen3.8-27B-Uncensored-OrcaRouter — MLX 5-bit

GGUF QUANTS HERE: [`chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-GGUF`](https://huggingface.co/chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-GGUF)

This unofficial native MLX affine 5-bit/group-64 checkpoint was independently generated directly from the pinned floating F16 parent orcarouter/Qwen3.8-27B-Uncensored-GGUF@402c3e0a64d77880f55ab096c5b7597ef85162ab. It was not transcoded from another quant and does not use the older FP8-expanded BF16 working representation. This conversion adds no training, fine-tuning, merging, or alignment change.

Source publisher notice — OrcaRouter says

[OrcaRouter claims:](https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored-GGUF/blob/402c3e0a64d77880f55ab096c5b7597ef85162ab/README.md)

An abliterated (refusal-removed) build of Qwen/Qwen3.8-27B, a dense native vision-language model with an MTP speculative-decoding head. ⚠️ Disclaimer — read before use This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). As a direct consequence: - It will comply with harmful, unethical, offensive, or illegal requests that the original Qwen3.8-27B would refuse. It has no meaningful built-in guardrails. - It is released strictly for legitimate research — interpretability, AI-safety and refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments. - You assume full responsibility and liability for how you use it and for everything it generates. Do not deploy it to end users or in production without adding your own safety, moderation, and abuse-prevention layers. - Use must comply with the Apache 2.0 License inherited from the base model, and all laws and regulations that apply to you. - The authors and uploaders accept no liability for any misuse or harm arising from this model. Its outputs do not reflect the views of the uploaders or of Qwen / Alibaba. By downloading or using this model you acknowledge and accept the above.

These are the source publisher's claims and warnings. This conversion has not independently revalidated refusal-removal, quality, safety, benchmark, or behavioral claims.

Direct floating source and precision audit

Source: orcarouter/Qwen3.8-27B-Uncensored-GGUF@402c3e0a64d77880f55ab096c5b7597ef85162ab, two F16 GGUF shards plus mmproj-Qwen3.8-27B-Uncensored-f16.gguf.

Sampled text and formerly quantized MTP matrices did not lie on the old FP8×scale lattice, establishing higher-precision pre-FP8 values in those samples. A dense mtp.fc.weight sample was exactly the old BF16 values stored as F16. A sampled vision-projector BF16 matrix was byte-identical to the old source-derived BF16 matrix despite the new projector's f16 filename. The projector's two distinct GGUF patch-embedding tensors are ordered temporal slices; both are stacked before the native MLX channels-last transpose, reconciling 334 physical mmproj tensors to 333 native vision tensors without dropping either slice. These audits establish precision lineage, not an intelligence gain.

Model tree

text
Qwen/Qwen3.8-27B
└── orcarouter/Qwen3.8-27B-Uncensored-GGUF@402c3e0a64d77880f55ab096c5b7597ef85162ab (F16 parent)
    └── chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-MLX-5bit

GGUF is the upstream container of the newly released floating weights; this child is a native MLX package reconstructed from those floating tensors.

MLX storage avoids an unnecessary pre-quantization loss: F16 tensors remain F16 (including the two unquantized token-embedding and MTP-fc tensors), BF16 tensors retain their raw BF16 bytes, and eligible F16/BF16 matrices are quantized directly in that source dtype. A source F32 dense tensor is stored as BF16 only when a full-value BF16-cast→F32 bit comparison reproduces every finite value exactly; otherwise it remains F32. The expected promoted tensors, including both temporal patch slices, are old BF16-lattice values and therefore round-trip losslessly. The recursively validated package contains 1199 logical tensors: 504 eligible matrices encoded as native MLX affine 5-bit/group-64 weights, and 695 dense tensors retained in their source-derived F16/BF16/F32 contract. The package contains 2207 physical tensors totaling 21,446,863,360 tensor-payload bytes. must be replaced with the final recursively validated counts. The required M4 Pro gates exercise this mixed F16/BF16/F32 package.

Integrity and release gates

Pinned source files

FileBytesSHA-256
Qwen3.8-27B-Uncensored-F16-00001-of-00002.gguf27,908,108,288578926d4e6d94281e95a48d8e154c4667061a669e9016f2aa2b15101ab4363dc
Qwen3.8-27B-Uncensored-F16-00002-of-00002.gguf26,749,625,920c15e78454caec46b19dae22fa6915a9e77b917e7cceffdc64c265801f4eaa1f1
mmproj-Qwen3.8-27B-Uncensored-f16.gguf931,145,984add205b7bfdb3f71f6da36b0a82aa20928dd829a920878c602628cdfbebc5288

Published model shards

FileBytesSHA-256
model-00001-of-00005.safetensors4,868,235,736da8940d2d1d3eb1f8a0013152783e804c86d18beba54cd60598d5686fd662430
model-00002-of-00005.safetensors4,887,209,9929251b9d11bed21fe2e75630ebc963a4d783be8566832520ac42a39fb0e049fc4
model-00003-of-00005.safetensors4,865,140,024af2c50f96477d5b1a12839ef36e75c63e993b7dfac52969a4b4f5e694657f635
model-00004-of-00005.safetensors4,883,581,784737752cd777d9dea18c9b541443862cc325516e149230aabab62b96243023a4a
model-00005-of-00005.safetensors1,942,977,248650868b1769a9df64af1148acb86dd2cf16773e55ebc7743b14c59a90bf059e7

Artifact manifest: /workspace/orcarouter-f16/work/release/MLX-5bit/ARTIFACT-MANIFEST.json — 4a948cee47dc6e365e26ddc74296932ece9bf837393f09cbe2467b84c9e65333.

M4 Pro runtime

Validated sequentially on Mac M4 Pro with 48 GiB unified memory, macOS 26.6.1, runtime mlx-serve-26.8.7|mlx-0.32.0|mlx-c-fba4470b8907|llama.cpp-b10034. Text determinism, chat-template tool calling, image-grounded vision (red), and clean unload all passed. MTP-off produced zero draft tokens in 10.150830s; MTP-on drafted 66 tokens and accepted 66 in 6.809538s. The matched output SHA-256 was 65966537023045093dda6a4bf49057afef35319d2f5170c68435d3330c8cec10 in both modes. Durations are bounded smoke-test evidence, not performance benchmarks.

Bounded matched fidelity

ModelPerplexityTokens
old-fp8-derived-Q4KM5.80342048
new-f16-derived-Q4KM5.79622048
old-fp8-derived-Q5KM5.79012048
new-f16-derived-Q5KM5.77742048
old-fp8-derived-Q6_K5.85862048
new-f16-derived-Q6_K5.80342048
old-fp8-derived-Q8_05.79582048
new-f16-derived-Q8_05.81952048

Fidelity report: /workspace/orcarouter-f16/work/fidelity-v2/FIDELITY-REPORT.json — 9774c9205abac6feb2ff8759627dc5b2087e703d7141f42a658432a92de02e0d.

Machine-readable release evidence

The JSON below is the fail-closed release contract. It intentionally excludes README/self-manifest hashes to avoid circular evidence.

<!-- RELEASE-EVIDENCE-BEGIN -->

json
{
  "artifacts": {
    "model-00001-of-00005.safetensors": {
      "sha256": "da8940d2d1d3eb1f8a0013152783e804c86d18beba54cd60598d5686fd662430",
      "size": 4868235736
    },
    "model-00002-of-00005.safetensors": {
      "sha256": "9251b9d11bed21fe2e75630ebc963a4d783be8566832520ac42a39fb0e049fc4",
      "size": 4887209992
    },
    "model-00003-of-00005.safetensors": {
      "sha256": "af2c50f96477d5b1a12839ef36e75c63e993b7dfac52969a4b4f5e694657f635",
      "size": 4865140024
    },
    "model-00004-of-00005.safetensors": {
      "sha256": "737752cd777d9dea18c9b541443862cc325516e149230aabab62b96243023a4a",
      "size": 4883581784
    },
    "model-00005-of-00005.safetensors": {
      "sha256": "650868b1769a9df64af1148acb86dd2cf16773e55ebc7743b14c59a90bf059e7",
      "size": 1942977248
    }
  },
  "fidelity": {
    "corpus": {
      "path": "wikitext-2-raw/wiki.test.raw",
      "sha256": "173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08",
      "size": 1290590
    },
    "report": {
      "path": "/workspace/orcarouter-f16/work/fidelity-v2/FIDELITY-REPORT.json",
      "sha256": "9774c9205abac6feb2ff8759627dc5b2087e703d7141f42a658432a92de02e0d",
      "size": 17333
    },
    "results": [
      {
        "label": "old-fp8-derived-Q4_K_M",
        "perplexity": 5.8034,
        "token_count": 2048
      },
      {
        "label": "new-f16-derived-Q4_K_M",
        "perplexity": 5.7962,
        "token_count": 2048
      },
      {
        "label": "old-fp8-derived-Q5_K_M",
        "perplexity": 5.7901,
        "token_count": 2048
      },
      {
        "label": "new-f16-derived-Q5_K_M",
        "perplexity": 5.7774,
        "token_count": 2048
      },
      {
        "label": "old-fp8-derived-Q6_K",
        "perplexity": 5.8586,
        "token_count": 2048
      },
      {
        "label": "new-f16-derived-Q6_K",
        "perplexity": 5.8034,
        "token_count": 2048
      },
      {
        "label": "old-fp8-derived-Q8_0",
        "perplexity": 5.7958,
        "token_count": 2048
      },
      {
        "label": "new-f16-derived-Q8_0",
        "perplexity": 5.8195,
        "token_count": 2048
      }
    ]
  },
  "kind": "mlx",
  "mlx": {
    "bits": 5,
    "bits_passed": [
      4,
      5,
      6,
      8
    ],
    "conversion_report": {
      "path": "/workspace/orcarouter-f16/work/state/mlx-conversion-runtime.json",
      "sha256": "649cd87ca10216f9743640de5f18ae7c42767efaca8e4e4f851fad5d46470092",
      "size": 1024
    },
    "dequantize_passed": true,
    "group_size": 64,
    "m4_runtime_passed": true,
    "mode": "affine",
    "quantize_passed": true
  },
  "runtime": {
    "gates": {
      "deterministic": {
        "passed": true,
        "run_1_output_sha256": "2f2d059139883d9985ea4178113bf6ac4fa731a1b72ab0be2608a3a279411144",
        "run_2_output_sha256": "2f2d059139883d9985ea4178113bf6ac4fa731a1b72ab0be2608a3a279411144"
      },
      "mtp_off": {
        "draft_tokens": 0,
        "mtp_weights_loaded": true,
        "passed": true,
        "speculative_decoding_enabled": false
      },
      "mtp_on": {
        "draft_tokens": 66,
        "mtp_weights_loaded": true,
        "passed": true,
        "speculative_decoding_enabled": true
      },
      "text": {
        "passed": true
      },
      "tool": {
        "passed": true
      },
      "vision": {
        "passed": true
      }
    },
    "hardware": {
      "macOS": "26.6.1",
      "memory_GiB": 48,
      "model": "Mac M4 Pro"
    },
    "report": {
      "path": "/workspace/orcarouter-f16/work/runtime-mlx/MLX-5bit/attestation.rebased.json",
      "sha256": "853e5b2ea3d1f5d03e39abf75549532ce5af6aaef44fb254c10b4bccdbdfc607",
      "size": 9869
    },
    "revision": "mlx-serve-26.8.7|mlx-0.32.0|mlx-c-fba4470b8907|llama.cpp-b10034"
  },
  "schema": 1,
  "source": {
    "files": [
      {
        "path": "Qwen3.8-27B-Uncensored-F16-00001-of-00002.gguf",
        "sha256": "578926d4e6d94281e95a48d8e154c4667061a669e9016f2aa2b15101ab4363dc",
        "size": 27908108288
      },
      {
        "path": "Qwen3.8-27B-Uncensored-F16-00002-of-00002.gguf",
        "sha256": "c15e78454caec46b19dae22fa6915a9e77b917e7cceffdc64c265801f4eaa1f1",
        "size": 26749625920
      },
      {
        "path": "mmproj-Qwen3.8-27B-Uncensored-f16.gguf",
        "sha256": "add205b7bfdb3f71f6da36b0a82aa20928dd829a920878c602628cdfbebc5288",
        "size": 931145984
      }
    ],
    "repo": "orcarouter/Qwen3.8-27B-Uncensored-GGUF",
    "revision": "402c3e0a64d77880f55ab096c5b7597ef85162ab"
  }
}

<!-- RELEASE-EVIDENCE-END -->

The bounded WikiText-2 comparison is quantization-fidelity evidence, not an intelligence benchmark. Throughput depends on hardware, context, prompt, and runtime. M4 Pro load/generation tests must cover MTP off and on before publication.

Legacy release history

The prior public revision was quantized from an FP8-expanded BF16 working representation. That historical lineage does not describe this replacement.