HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8
Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8
Reduced-refusal model. A refusal-direction edit was deliberately applied first in BF16, then the edited BF16 was quantized with NVIDIA ModelOpt. See the safety caveat below. Darkstar is the HangGlidersRule tuning brand. Internal edit-lineage id: R3. Prior llm-compressor compressed-tensors NVFP4 builds are rejected/historical. This build reuses the selected clean-base mixed W4A16-NVFP4+FP8 recipe and is independent of, and not blocked on, any other product.
Private checkpoint repository: HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8.
This card describes the selected (promoted) Abliterated ModelOpt product, and its id encodes that artifact's real precision class. The uniform Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A4-NVFP4 id stays reserved and unbuilt. A bare Abliterated-NVFP4 target is never reserved or published, and Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-NVFP4 remains the abstract release slot in the ledger only.
Summary
ModelOpt NVFP4 quantization of the Darkstar Abliterated BF16 derivative (internal lineage R3). The refusal-direction edit is applied first in BF16; NVFP4 quantization follows with the selected clean mixed W4A16-NVFP4+FP8 recipe. The mixed candidate is the finished local build.
Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-NVFP4 is the family release slot. Its concrete candidates are named for their real precision class and mirror the clean-base candidates:
- `Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8` — the selected (promoted) Abliterated ModelOpt product. Mixed precision (W4A16 NVFP4 on language MLP +
lm_head, FP8 on self-attention and GatedDeltaNet projections, BF16 protected/KV). Its single-stream winner is MTP10 at mean 251.889 tok/s, with GPQA148/198 = 74.75%. - `Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A4-NVFP4` — uniform W4A4 NVFP4. Not built (the clean-base W4A4 comparison already rejected uniform W4A4 on throughput).
Candidate precision maps
<!-- CANDIDATE-SYNC candidate=Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8 -->
<!-- CANDIDATE-SYNC candidate=Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A4-NVFP4 -->
Provenance
- Upstream model:
Qwen/Qwen3.8-27B - Upstream revision:
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 - Editable source:
HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-BF16(see `bf16.md`) - Weight edit: refusal-direction projection (abliteration), layer 38, seed 42
- Exact recipe: `recipes/qwen3.8-27b/darkstar-qwen3.8-27b-abliterated-modelopt-nvfp4.yaml`
- Exact selected operator recipe (tracked): `configs/modelopt/recipes/w4a16_nvfp4_mse-fp8_attn-kv_bf16.yaml`, SHA-256
90fc6b37c00334debd49f1975ab406b5e20667f07e4be0be3e463a648abac642— the identical selected clean-base recipe. It quantizeslm_headin W4A16 NVFP4, matching the product precision map above. - Artifact identity:
_SUCCESS.jsonSHA-2563d89ec57c1371e142adc2584de079b54a0e1d8c12dc9550118d0a851da020a79;manifest.sha256SHA-256642dbbe89b085a2daf5119c37c0496576a475ed64c36653fc993c04abaf2ca9f - ModelOpt pin/recipes: `modelopt/README.md`
- Engineering repository: `HangGlidersRule/model-forge`
- Lineage detail: `artifact-lineage.md`
Edit + quantization summary
- Edit: normalized float32 refusal-direction projection at layer 38, seed 42, applied to exactly 131 residual-writing tensors (see `bf16.md` for the full inventory). Vision tower untouched.
- Quantization: NVIDIA ModelOpt unified HF NVFP4/FP8 (
hf_quant_config.json), reusing the selected clean-base mixed W4A16-NVFP4+FP8 recipe (W4A16 NVFP4 on language MLP +lm_head, FP8 on self-attention and GatedDeltaNet projections). Prior compressed-tensors builds are rejected/historical. - Kept BF16 in the mixed candidate: vision, MTP,
conv1d, norms, embeddings, and runtime KV. - Runtime KV: BF16 (no FP8 KV metadata).
- Calibration:
cnn_dailymail+nemotron-post-training-dataset-v2, 512+512, seq 2048, seed 1234. - MTP: preserve/reattach all 15 source BF16 MTP tensors via ModelOpt's MTP path.
- Validation: fail-closed validators passed (no NaN/Inf/zero scales, no quantized vision, 15 BF16 MTP tensors, no FP8 KV metadata, no mixed fused groups); 131 R3-edited tensors preserved through quant.
Intended use
- Reduced-refusal, high-throughput text generation on Blackwell-class GPUs.
- Research on the combined effect of the refusal edit and NVFP4 quantization.
Safety caveat (reduced refusal)
This model has had its refusal direction deliberately reduced in the BF16 source. It will comply with many requests upstream would refuse and has no added safety mitigations. A fresh harmful-refusal and safe-over-refusal evaluation was measured directly on this quantized build (283/283 terminal: 200/200 harmful compliance, 0/83 safe over-refusals, 0 errors). Deploy only behind your own policy, filtering, and access controls, and only where lawful. These behavior numbers are measurements, not safety endorsements.
Limitations
- GPQA Diamond for this matched thinking-off NVFP4 cell is measured on the selected mixed
W4A16-NVFP4-Mixed-FP8candidate: 148/198 = 74.75%, full denominator (198/198 terminal parseable, 0 timeout/parse/error). A secondary thinking-enabled run scored164/198 = 82.83%, but that number is from the rejected historical R3 compressed-tensors NVFP4 artifact, not this ModelOpt build; it is not matched-matrix eligible and cannot be attributed purely to quantization. - The R3 edit itself costs measured accuracy (full-denominator BF16 delta
-11questions /-5.56pp; see `bf16.md`). - Quantization can introduce quality regressions not captured by the smoke suite.
Evaluation
Curated aggregates: `../results/gpqa-matrix.json`. Protocol: `../gpqa-protocol.md`. Full caveats: `../benchmark-matrix.md`.
Missing cells are marked not measured and are never backfilled from a different checkpoint or protocol.
Full-denominator, measured on this exact build. The GPQA row above is a verified full-denominator measurement on the selected mixed candidate (thinking off, temperature 1.0, top-p 0.95, top-k 20, 4 workers, no output cap; external operator evidence with immutable hashes in `../results/gpqa-matrix.json`). GPQA question text, answer keys, per-question responses, and the run journal are intentionally not committed.
Runtime requirements and example
- Runtime: vLLM
0.27.1, compiled mode, Flash Attention, BF16 KV cache. - Context length: up to 126,144.
- Serve profile: native MTP; the selected candidate's performance winner is MTP depth 10, 32K scheduler budget,
max_num_seqs=16, prefix caching + chunked prefill. The MTP1-12 sweep was non-monotonic (headline peak MTP8 251.316 tok/s), so MTP10 was selected on its higher mean throughput (251.889 tok/s vs MTP8 mean 250.862 tok/s). - Serving correctness: tools, strict JSON, vision, prefix cache, and 20-request sustained load all pass; evidence SHA-256
4c88632efcfc736518a66351c95735eb2f9ff7ce79496d050a53d171beaf4613. - API alias:
darkstar-qwen38-abliterated-nvfp4; container:vllm-darkstar-qwen38-abliterated-modelopt.
VLLM_ATTENTION_BACKEND=FLASH_ATTN \
vllm serve HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8 \
--served-model-name darkstar-qwen38-abliterated-nvfp4 \
--kv-cache-dtype bf16 \
--max-model-len 126144 \
--max-num-seqs 16 \
--max-num-batched-tokens 32768 \
--enable-chunked-prefill \
--enable-prefix-caching \
--compilation-config 2 \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 10}'See the exact frozen profile in `containers/serve/darkstar-qwen38-abliterated-nvfp4.yml`. The checked-in canonical launcher supports --dry-run and --print-config, validates the deterministic tracked Compose digest, and allows only the host port as a Product 4 environment override. It does not accept mutable vLLM model/runtime arguments.
Serving capacity (per-concurrency, measured on this build)
Matched capacity cells at the frozen MTP10 profile: 512 generated tokens per request, two repeats per cell, concurrency 1 and 2 — the only concurrency levels this run measured, so no higher concurrency is reported. Zero failed requests and zero fatal markers in every cell. Machine-readable record: `../results/serving-capacity-profiles.json`; source evidence SHA-256 6e52a5ad4f87a8b12866e0939c2d2024701172d8b5a56a7839ce00738f1a3ac9.
Concurrency capacity is reported separately from the single-stream throughput winner above and is never mixed into it.
Publication-readiness (rendered from the ledger)
This block is rendered from the machine-readable source of truth `../results/publication-readiness-ledger.json` and kept in sync by CI, per the four-product release process. The benchmark matrix shows all four products together. This checkpoint is public on Hugging Face with clean download/boot/smoke verified. The immutable Git tag is the sole remaining release gate.
<!-- LEDGER-SYNC product=Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-NVFP4 -->
License and attribution
- License: Apache-2.0.
- Preserve upstream attribution and required notices.
Serving-stack update (2026-09-16): NVFP4-KV production profile + measured gates
A production serve profile update was measured on this exact artifact (2026-09-16, mcprue RTX PRO 6000 Blackwell, single GPU). It is an alternate serving profile — the artifact weights are unchanged (frozen at _SUCCESS SHA-256 3d89ec57...).
NVFP4-KV serving profile (validated 2026-09-16)
- Runtime: custom vLLM nightly
0.27.2rc1.dev77+gac7509e2bwith upstream PR #49891 (FA2-nvfp4 routing) rebased + an sm120 linear-V-scale writer overlay (imagevllm-qwen38:nvfp4kv). - KV cache: nvfp4 (4-bit KV) with
--attention-backend FLASHINFER(top + spec config). - Spec decode: MTP depth 4 (
{"method":"mtp","num_speculative_tokens":4,"attention_backend": "FLASHINFER"}), mamba cache modealign, chunked prefill. - Context: 262,144 (native
max_position_embeddings), util 0.50, mnbt 8192, seqs 16. - KV pool: 1,046,192 tokens at util 0.50 (~4× the bf16-KV pool at equal VRAM budget); max concurrency 3.99× at full 262,144-token requests.
Measured gates on the NVFP4-KV profile (secondary protocol — llm-inference-bench,
thinking-on, concurrency 30, max-tokens 32768; NOT the matched-matrix cell above)
- GPQA Diamond: 170/198 = 85.86% (0 errors, 0 unparseable, 13 truncated, 0 IMAs). Reference runs on the same protocol: bf16-KV serve 169/198 = 85.35%; earlier nvfp4-KV boot 168/198 = 84.85%. The quantized-KV profile is parity-or-better with the bf16-KV serving reference while carrying 4× the KV pool. Protocol mismatch note: these numbers use a different harness/protocol than the matched thinking-off cell (148/198) above and are not comparable to it; they compare serving profiles to each other only.
- Long-context needles: HIT at 52,080 / 100,079 / 145,078 / 210,078 prompt tokens.
- Sustained MTP acceptance under thinking-on load: ~63-70% of drafted tokens accepted.
- Zero illegal-memory-access events across the full concurrency-30 thinking-load battery (the fp8-KV Triton path failed this battery; the nvfp4 path passes it).
Behavior-gate band disclosure (decision A, 2026-09-16)
The gate re-measured on modern serving stacks (7 runs, canonical harness, raw completions, temp 0, thinking-independent) gives a band: 197-198/200 harmful compliance + 0-1/83 over-refusals, with borderline items (suicide-adjacent crisis preambles; a legal-info clarification) flipping between boots — boot-level greedy tie-break nondeterminism at the decision edge. The 200/200 + 0/83 figures above were single-boot measurements in the release window (the exact 2026-08-21 release-gate serve configuration reproduces 197-198/200 + 0/83 today; MTP10 + xxhash config restores 0/83 — MTP4 causes one over-refusal). Serving stack, KV dtype, and runtime version are ruled out (identical scores across all configurations). Honest current claim: ~98.5-99% harmful compliance band, 0-1/83 over-refusal, not a stable 200/200.
