Marcin116/Qwen3.8-27B-Uncensored-W4A16-AutoRound-embed-int4
Qwen3.8-27B-Uncensored-W4A16-AutoRound-embed-int4
A complete, single-repository W4A16 checkpoint of `orcarouter/Qwen3.8-27B-Uncensored` (Apache-2.0, abliterated Qwen3.8-27B). The transformer body is symmetric INT4 / group-128 compressed-tensors (pack-quantized) — the same recipe as `born2bewild/Qwen3.8-27B-W4A16-AutoRound-fast` — and the token-embedding group (model.language_model.embed_tokens) is INT4 (not the INT8 used in that source). Everything else follows the orcarouter Qwen3.8 architecture unchanged, which unlike the born2bewild source has no mtp.draft_lm_head ( mtp_use_dedicated_embeddings: false): the single mtp. module reuses the shared lm_head for its output projection.
[!WARNING] This checkpoint inherits the substantial safety-alignment removal of the abliterated upstream model. It can comply with harmful, unethical, offensive, or illegal requests and should not be treated as having meaningful built-in guardrails. It is intended for legitimate local use, evaluation, red-teaming, and safety research. Do not expose it to untrusted users without an independent safety layer.
Provenance
- `Qwen/Qwen3.8-27B` — original Apache-2.0 model.
- `orcarouter/Qwen3.8-27B-Uncensored` — abliterated derivative, BF16, 18 shards.
- This checkpoint — W4A16 quantization built directly from (2).
The quantization recipe is the M1 recipe (born2bewild/Qwen3.8-27B-W4A16-AutoRound-fast) applied to the uncensored base, minus the MTP draft head that does not exist in this architecture. The `Marcin116/Qwen3.8-27B-Uncensored-W4A16-RTX3090-MTP4-embed-int4` repository is the sibling "fast" build (GPTQ-calibrated heads, draft vocabulary, RTX-3090 serving overlay); this repository is the plain AutoRound-style W4A16 build of the same uncensored weights.
Quantization layout
Symmetric INT4, group size 128, pack-quantized, compressed-tensors, quantization_status: compressed (quant_method: compressed-tensors). Per group: scale = clamp(absmax/7, min=1e-10), codes clamp(round(x/scale), -7, +7) packed 8-per-INT32; BF16 scale per 128-element group; weight_shape (I64 [2]) stores the logical shape.
Quantized to INT4 (410 modules):
Kept in BF16 (789 tensors): all layernorms (incl. model.language_model.norm), GDN recurrence-control tensors (linear_attn.{in_proj_a,in_proj_b,conv1d,dt_bias,A_log,norm} × 48 layers), the entire vision tower model.visual.* (333 tensors), and the MTP norms (mtp.{norm,pre_fc_norm_embedding,pre_fc_norm_hidden}, mtp.layers.0.{input_layernorm,post_attention_layernorm,self_attn.{q_norm,k_norm}}).
The layer stack is 48 GatedDeltaNet layers interleaved with 16 full-attention layers (1 in every 4). The config.json ignore list holds the 207 module names that stay BF16 (111 vision modules — 2-D projections plus pos_embed — and 96 GDN in_proj_a/b); the standalone quantization_config.json holds the simplified group_0-only schema with 208 entries (those 207 + lm_head).
Token embeddings: INT4 (this variant)
The embeddings were quantized from the uncensored BF16 ground truth (the abliterated orcarouter/Qwen3.8-27B-Uncensored embed_tokens), not from the base Qwen embeddings — the abliterated model carries fine-tuned embeddings, so its BF16 weights are the correct reference.
Same pack-quantized group-128 symmetric layout as every other INT4 group; the loader reconstructs the full [248320, 5120] matrix from weight_shape. config.json marks it: group_2.targets = ["re:.*embed_tokens$"], weights = {type: int, symmetric: true, group_size: 128, num_bits: 4}.
Size
model.safetensors.index.json declares total_size = 15,047,947,146 bytes (= exact on-disk sum of the 18 shards) and total_parameters = 27,781,427,952 (sum of logical pre-quant tensor sizes, incl. vision tower and MTP). Vocabulary 248,320, hidden 5,120, 64 language layers; mtp_use_dedicated_embeddings: false, no mtp.draft_lm_head.
There are no model_extra_tensors.safetensors / mtp_draft_vocab_ids.pt files: this variant carries the native 1-layer MTP only and uses the shared lm_head, so no extra tensors are needed for from_pretrained().
Serving
Compatible with vLLM (or any runtime with Qwen3.8 compressed-tensors / pack-quantized INT4 support). Typical launch:
vllm serve <this-repo> \
--max-model-len 32768 \
--gpu-memory-utilization 0.85For MTP speculative decoding with the native MTP module, enable the runtime's MTP spec path (e.g. vLLM --speculative-config '{"method": "mtp", "num_speculative_tokens": N}') as supported by the runtime version; MTP acceptance depends on the workload and affects speed, not quality.
Quality / safety notes
- The W4A16 body follows the standard AutoRound-style symmetric INT4 group-128 recipe used for this model family; no per-model calibration data is shipped with this repository.
- The INT4 token embedding is a single-tensor reduction from the BF16 uncensored ground truth: per-group symmetric 4-bit quantization (
absmax/7scaling). Measured against the BF16 ground truth, its per-element error is in the same band as the INT4 linear layers (11.8–12.5% rel-RMS), not below it. The difference is how often that error is invoked: an embedding is looked up once per token, whereas every INT4 linear is applied once per layer, so the embedding's quantization error has only one propagation step per token and its practical impact on output quality is minor. - Abliteration is inherited from the upstream: expect minimal refusal and no built-in safety layer.
Files
model-0000X-of-00018.safetensors— 18 model shards (same boundaries as the BF16 base; quantized tensors replaced in place by<module>.weight_packed/.weight_scale/.weight_shape).model.safetensors.index.json— index over all 2,019 post-quant keys.config.json— orcarouterqwen3_5config,dtype: bfloat16, in-configquantization_config(4 groups, INT4 everywhere, 207-entry ignore list).quantization_config.json— standalone simplified schema (quant_method: compressed-tensors,format: pack-quantized,quantization_status: compressed,group_0Linear INT4, 208-entry ignore list includinglm_head).tokenizer.json,tokenizer_config.json,merges.txt,vocab.json,chat_template.jinja,generation_config.json,preprocessor_config.json,video_preprocessor_config.json,LICENSE.
License and attribution
Apache-2.0, inherited from Qwen and the intermediate checkpoints. See `LICENSE`. This repository is a local derivative produced by Marcin116 from the upstream lineage above; the upstream repositories remain untouched.
