CoolFace
Modelpublic

HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
4likes3kdownloads
Model Card

Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8

Reduced-refusal model. A refusal-direction edit was deliberately applied first in BF16, then the edited BF16 was quantized with NVIDIA ModelOpt. See the safety caveat below. Darkstar is the HangGlidersRule tuning brand. Internal edit-lineage id: R3. Prior llm-compressor compressed-tensors NVFP4 builds are rejected/historical. This build reuses the selected clean-base mixed W4A16-NVFP4+FP8 recipe and is independent of, and not blocked on, any other product.

Private checkpoint repository: HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8.

This card describes the selected (promoted) Abliterated ModelOpt product, and its id encodes that artifact's real precision class. The uniform Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A4-NVFP4 id stays reserved and unbuilt. A bare Abliterated-NVFP4 target is never reserved or published, and Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-NVFP4 remains the abstract release slot in the ledger only.

Summary

ModelOpt NVFP4 quantization of the Darkstar Abliterated BF16 derivative (internal lineage R3). The refusal-direction edit is applied first in BF16; NVFP4 quantization follows with the selected clean mixed W4A16-NVFP4+FP8 recipe. The mixed candidate is the finished local build.

Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-NVFP4 is the family release slot. Its concrete candidates are named for their real precision class and mirror the clean-base candidates:

  • —`Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8` — the selected (promoted) Abliterated ModelOpt product. Mixed precision (W4A16 NVFP4 on language MLP + lm_head, FP8 on self-attention and GatedDeltaNet projections, BF16 protected/KV). Its single-stream winner is MTP10 at mean 251.889 tok/s, with GPQA 148/198 = 74.75%.
  • —`Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A4-NVFP4` — uniform W4A4 NVFP4. Not built (the clean-base W4A4 comparison already rejected uniform W4A4 on throughput).

Candidate precision maps

<!-- CANDIDATE-SYNC candidate=Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8 -->

ComponentPrecision
language_mlpW4A16-NVFP4-g16
lm_headW4A16-NVFP4-g16
self_attentionFP8-e4m3
gdn_projectionsFP8-e4m3
kv_cacheBF16
protectedBF16

<!-- CANDIDATE-SYNC candidate=Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A4-NVFP4 -->

ComponentPrecision
language_mlpW4A4-NVFP4-g16
lm_headBF16
self_attentionW4A4-NVFP4-g16
gdn_projectionsW4A4-NVFP4-g16
kv_cacheBF16
protectedBF16

Provenance

  • —Upstream model: Qwen/Qwen3.8-27B
  • —Upstream revision: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
  • —Editable source: HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-BF16 (see `bf16.md`)
  • —Weight edit: refusal-direction projection (abliteration), layer 38, seed 42
  • —Exact recipe: `recipes/qwen3.8-27b/darkstar-qwen3.8-27b-abliterated-modelopt-nvfp4.yaml`
  • —Exact selected operator recipe (tracked): `configs/modelopt/recipes/w4a16_nvfp4_mse-fp8_attn-kv_bf16.yaml`, SHA-256 90fc6b37c00334debd49f1975ab406b5e20667f07e4be0be3e463a648abac642 — the identical selected clean-base recipe. It quantizes lm_head in W4A16 NVFP4, matching the product precision map above.
  • —Artifact identity: _SUCCESS.json SHA-256 3d89ec57c1371e142adc2584de079b54a0e1d8c12dc9550118d0a851da020a79; manifest.sha256 SHA-256 642dbbe89b085a2daf5119c37c0496576a475ed64c36653fc993c04abaf2ca9f
  • —ModelOpt pin/recipes: `modelopt/README.md`
  • —Engineering repository: `HangGlidersRule/model-forge`
  • —Lineage detail: `artifact-lineage.md`

Edit + quantization summary

  • —Edit: normalized float32 refusal-direction projection at layer 38, seed 42, applied to exactly 131 residual-writing tensors (see `bf16.md` for the full inventory). Vision tower untouched.
  • —Quantization: NVIDIA ModelOpt unified HF NVFP4/FP8 (hf_quant_config.json), reusing the selected clean-base mixed W4A16-NVFP4+FP8 recipe (W4A16 NVFP4 on language MLP + lm_head, FP8 on self-attention and GatedDeltaNet projections). Prior compressed-tensors builds are rejected/historical.
  • —Kept BF16 in the mixed candidate: vision, MTP, conv1d, norms, embeddings, and runtime KV.
  • —Runtime KV: BF16 (no FP8 KV metadata).
  • —Calibration: cnn_dailymail + nemotron-post-training-dataset-v2, 512+512, seq 2048, seed 1234.
  • —MTP: preserve/reattach all 15 source BF16 MTP tensors via ModelOpt's MTP path.
  • —Validation: fail-closed validators passed (no NaN/Inf/zero scales, no quantized vision, 15 BF16 MTP tensors, no FP8 KV metadata, no mixed fused groups); 131 R3-edited tensors preserved through quant.

Intended use

  • —Reduced-refusal, high-throughput text generation on Blackwell-class GPUs.
  • —Research on the combined effect of the refusal edit and NVFP4 quantization.

Safety caveat (reduced refusal)

This model has had its refusal direction deliberately reduced in the BF16 source. It will comply with many requests upstream would refuse and has no added safety mitigations. A fresh harmful-refusal and safe-over-refusal evaluation was measured directly on this quantized build (283/283 terminal: 200/200 harmful compliance, 0/83 safe over-refusals, 0 errors). Deploy only behind your own policy, filtering, and access controls, and only where lawful. These behavior numbers are measurements, not safety endorsements.

Limitations

  • —GPQA Diamond for this matched thinking-off NVFP4 cell is measured on the selected mixed W4A16-NVFP4-Mixed-FP8 candidate: 148/198 = 74.75%, full denominator (198/198 terminal parseable, 0 timeout/parse/error). A secondary thinking-enabled run scored 164/198 = 82.83%, but that number is from the rejected historical R3 compressed-tensors NVFP4 artifact, not this ModelOpt build; it is not matched-matrix eligible and cannot be attributed purely to quantization.
  • —The R3 edit itself costs measured accuracy (full-denominator BF16 delta -11 questions / -5.56 pp; see `bf16.md`).
  • —Quantization can introduce quality regressions not captured by the smoke suite.

Evaluation

Curated aggregates: `../results/gpqa-matrix.json`. Protocol: `../gpqa-protocol.md`. Full caveats: `../benchmark-matrix.md`.

MetricValueBasis
GPQA Diamond (thinking off, matched)148/198 = 74.75%198/198 terminal parseable; 0 timeout/parse/error; summary SHA d8d0b5c0de686846338ce89e9a55456baec0550bbad765ccc65e9fa57380b818; journal SHA 9bb4913202977bad204ebde8d2e31e8357a3308f3b77c44539ef3977a2c6e813
GPQA Diamond (thinking on, secondary)164/198 = 82.83%rejected historical R3 compressed-tensors NVFP4; not matched-matrix eligible
Quantization delta vs Abliterated BF16+2 questions / +1.01 ppAbliterated BF16 146/198 → mixed 148/198, full denominator on both
Harmful-prompt compliance200/200 (0/200 refusals)283/283 terminal; 0 errors; summary SHA d814eac6eef86cb32c891d5c3b1765be806cb0fb634173080cd5df46ea9f9233; journal SHA 7b6ddf556ab3afc1f8582041d7b723dbfeaeb24b02b8bd27562b0c9928a37d4f
Safe over-refusals0/83 (0.00%)over-refusal suite on this exact ModelOpt build, 0 errors
Single-stream throughput (MTP10)mean 251.889 tok/snonmonotonic MTP1-12 sweep; MTP8 headline peak 251.316 tok/s; confirmation 10->8->8->10 selected MTP10 over MTP8 mean 250.862; confirmation SHA 6e52a5ad4f87a8b12866e0939c2d2024701172d8b5a56a7839ce00738f1a3ac9

Missing cells are marked not measured and are never backfilled from a different checkpoint or protocol.

Full-denominator, measured on this exact build. The GPQA row above is a verified full-denominator measurement on the selected mixed candidate (thinking off, temperature 1.0, top-p 0.95, top-k 20, 4 workers, no output cap; external operator evidence with immutable hashes in `../results/gpqa-matrix.json`). GPQA question text, answer keys, per-question responses, and the run journal are intentionally not committed.

Runtime requirements and example

  • —Runtime: vLLM 0.27.1, compiled mode, Flash Attention, BF16 KV cache.
  • —Context length: up to 126,144.
  • —Serve profile: native MTP; the selected candidate's performance winner is MTP depth 10, 32K scheduler budget, max_num_seqs=16, prefix caching + chunked prefill. The MTP1-12 sweep was non-monotonic (headline peak MTP8 251.316 tok/s), so MTP10 was selected on its higher mean throughput (251.889 tok/s vs MTP8 mean 250.862 tok/s).
  • —Serving correctness: tools, strict JSON, vision, prefix cache, and 20-request sustained load all pass; evidence SHA-256 4c88632efcfc736518a66351c95735eb2f9ff7ce79496d050a53d171beaf4613.
  • —API alias: darkstar-qwen38-abliterated-nvfp4; container: vllm-darkstar-qwen38-abliterated-modelopt.
bash
VLLM_ATTENTION_BACKEND=FLASH_ATTN \
vllm serve HangGlidersRule/Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-W4A16-NVFP4-Mixed-FP8 \
  --served-model-name darkstar-qwen38-abliterated-nvfp4 \
  --kv-cache-dtype bf16 \
  --max-model-len 126144 \
  --max-num-seqs 16 \
  --max-num-batched-tokens 32768 \
  --enable-chunked-prefill \
  --enable-prefix-caching \
  --compilation-config 2 \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 10}'

See the exact frozen profile in `containers/serve/darkstar-qwen38-abliterated-nvfp4.yml`. The checked-in canonical launcher supports --dry-run and --print-config, validates the deterministic tracked Compose digest, and allows only the host port as a Product 4 environment override. It does not accept mutable vLLM model/runtime arguments.

Serving capacity (per-concurrency, measured on this build)

Matched capacity cells at the frozen MTP10 profile: 512 generated tokens per request, two repeats per cell, concurrency 1 and 2 — the only concurrency levels this run measured, so no higher concurrency is reported. Zero failed requests and zero fatal markers in every cell. Machine-readable record: `../results/serving-capacity-profiles.json`; source evidence SHA-256 6e52a5ad4f87a8b12866e0939c2d2024701172d8b5a56a7839ce00738f1a3ac9.

PromptPrompt tokensC1 mean aggregate tok/s (pass 1 / pass 4)C2 mean aggregate tok/s (pass 1 / pass 4)
4K chars738193.411 / 195.486352.313 / 356.564
16K chars2653183.611 / 183.134328.393 / 343.799
48K chars7758157.579 / 157.889286.029 / 284.465

Concurrency capacity is reported separately from the single-stream throughput winner above and is never mixed into it.

Publication-readiness (rendered from the ledger)

This block is rendered from the machine-readable source of truth `../results/publication-readiness-ledger.json` and kept in sync by CI, per the four-product release process. The benchmark matrix shows all four products together. This checkpoint is public on Hugging Face with clean download/boot/smoke verified. The immutable Git tag is the sole remaining release gate.

<!-- LEDGER-SYNC product=Darkstar-Qwen3.8-27B-Abliterated-ModelOpt-NVFP4 -->

GateStatusEvidenceNext required proof
provenance_ownershipverifiedSource pinned to Qwen/Qwen3.8-27B@1d4bf0f...; R3 refusal-direction edit lineage documented; Darkstar-owned ModelOpt NVFP4 quantizationnone — gate satisfied
artifact_manifestverifiedLocal build frozen at ${PUBLICARTIFACTPATH}: _SUCCESS SHA-256 3d89ec57..., manifest.sha256 SHA-256 642dbbe8...; runtime snapshot inspect prefix f1e94a76, logs 4ffe8666, operator snapshot compose 85ba6815 recorded; tracked serve compose 5434c2a9 rendered from the frozen profilenone — gate satisfied
recipeeditmanifestverifiedPinned recipe present; reuses the selected clean mixed W4A16-NVFP4+FP8 recipe (recipe_sha256 90fc6b37...) plus the 131-tensor R3 edit inventorynone — gate satisfied
artifact_validationverifiedFail-closed validators passed on the export: no NaN/Inf/zero scales, no quantized vision, 15 BF16 MTP tensors, no FP8 KV metadata, no mixed fused groups, 131 R3-edited tensors preserved through quantization; snapshot capturednone — gate satisfied
abliteration_passverifiedReproducible projection (layer 38, seed 42, leakage ~0.000195); fresh abliteration eval on this exact ModelOpt build: 200/200 harmful compliance (0/200 refusals), 0/83 over-refusals, 0 errors; summary SHA d814ea..., journal SHA 7b6dd...none — gate satisfied
modeloptcandidatecomparisonverifiedReuses the selected clean-base W4A16-vs-W4A4 comparison; the abliterated mixed W4A16-NVFP4+FP8 candidate is built and measured (single-stream winner MTP10 251.889 tok/s); uniform W4A4 abliterated not built (clean-base W4A4 already rejected on throughput); winning recipe recordednone — gate satisfied
nvfp4tensorscale_validationverifiedTensor/scale integrity validated on the abliterated mixed build: finite NVFP4 scales, BF16 MTP preserved, no quantized vision, no mixed fused groups, no FP8 KV metadatanone — gate satisfied
performance_profileverifiedIndependent single-stream MTP1-12 sweep on the abliterated mixed candidate: headline peak MTP8 251.316 tok/s, but non-monotonic; confirmation 10->8->8->10 selected MTP10 (mean 251.889 tok/s vs MTP8 mean 250.862 tok/s), 32K scheduler budget, max-num-seqs 16, context 126144, BF16 KV, FlashAttentionnone — gate satisfied
servingcapacityprofileverifiedMatched per-concurrency cells measured on this exact build at the selected MTP10 profile: 4K/16K/48K prompts at concurrency 1 and 2, 512 generated tokens, 2 repeats per cell, 0 failed requests, zero fatal markers (source SHA 6e52a5ad...); semantic serving gates (tools, strict JSON, vision, prefix cache, 20-request sustained load) all pass, evidence SHA 4c8863...none — gate satisfied
gpqamatchedfull_denominatorverifiedAbliterated mixed W4A16-NVFP4+FP8 candidate: 148/198 = 74.75% full denominator, 198/198 terminal parseable, 0 timeout/parse/error, thinking off, temp 1.0 topp 0.95 topk 20; external operator evidence with immutable hashes (summary d8d0b5c0..., journal 9bb49132...)none — gate satisfied
behaviorrefusalevalverifiedFresh abliteration eval on this exact ModelOpt build: 283/283 terminal (200 harmful + 83 safe), 200/200 harmful compliance (0/200 refusals), 0/83 safe over-refusals, 0 errorsnone — gate satisfied
serveprofilefrozenverifiedFrozen serve profile: vLLM compiled mode, FlashAttention, BF16 KV, context 126144, MTP depth 10, 32K scheduler budget, max-num-seqs 16, prefix caching + chunked prefillnone — gate satisfied
modelcardfinalverifiedFinal model card is complete and references planned immutable release tag darkstar-qwen3.8-27b-v1.0.0; no placeholders remainnone — gate satisfied
publicationtargetshf_ghcrverifiedAnonymous Hugging Face API and config.json checks verify the repository is public, ungated, enabled, and at revision 2e25bd97fd1b6e6c7989e74c261d93a8702496e8; GHCR is explicitly not required for this releasenone — gate satisfied
cleandownloadboot_smokeverifiedFresh-target download verified with zero failures; vLLM boot and models/text/strict JSON/tool/vision smoke passed with zero failures and an empty fatal log; public revision 2e25bd97fd1b6e6c7989e74c261d93a8702496e8none — gate satisfied
release_tagverifiedImmutable Git release tag darkstar-qwen3.8-27b-v1.0.0 exists and is referenced from the final model cardnone — gate satisfied
noinheritedunverified_resultsverifiedGPQA, throughput, serving correctness, and the fresh abliteration eval all measured directly on this exact abliterated ModelOpt build; compressed-tensors GPQA/perf/behavior numbers rejected, never inheritednone — gate satisfied

License and attribution

  • —License: Apache-2.0.
  • —Preserve upstream attribution and required notices.

Serving-stack update (2026-09-16): NVFP4-KV production profile + measured gates

A production serve profile update was measured on this exact artifact (2026-09-16, mcprue RTX PRO 6000 Blackwell, single GPU). It is an alternate serving profile — the artifact weights are unchanged (frozen at _SUCCESS SHA-256 3d89ec57...).

NVFP4-KV serving profile (validated 2026-09-16)

  • —Runtime: custom vLLM nightly 0.27.2rc1.dev77+gac7509e2b with upstream PR #49891 (FA2-nvfp4 routing) rebased + an sm120 linear-V-scale writer overlay (image vllm-qwen38:nvfp4kv).
  • —KV cache: nvfp4 (4-bit KV) with --attention-backend FLASHINFER (top + spec config).
  • —Spec decode: MTP depth 4 ({"method":"mtp","num_speculative_tokens":4,"attention_backend": "FLASHINFER"}), mamba cache mode align, chunked prefill.
  • —Context: 262,144 (native max_position_embeddings), util 0.50, mnbt 8192, seqs 16.
  • —KV pool: 1,046,192 tokens at util 0.50 (~4× the bf16-KV pool at equal VRAM budget); max concurrency 3.99× at full 262,144-token requests.

Measured gates on the NVFP4-KV profile (secondary protocol — llm-inference-bench,

thinking-on, concurrency 30, max-tokens 32768; NOT the matched-matrix cell above)

  • —GPQA Diamond: 170/198 = 85.86% (0 errors, 0 unparseable, 13 truncated, 0 IMAs). Reference runs on the same protocol: bf16-KV serve 169/198 = 85.35%; earlier nvfp4-KV boot 168/198 = 84.85%. The quantized-KV profile is parity-or-better with the bf16-KV serving reference while carrying 4× the KV pool. Protocol mismatch note: these numbers use a different harness/protocol than the matched thinking-off cell (148/198) above and are not comparable to it; they compare serving profiles to each other only.
  • —Long-context needles: HIT at 52,080 / 100,079 / 145,078 / 210,078 prompt tokens.
  • —Sustained MTP acceptance under thinking-on load: ~63-70% of drafted tokens accepted.
  • —Zero illegal-memory-access events across the full concurrency-30 thinking-load battery (the fp8-KV Triton path failed this battery; the nvfp4 path passes it).

Behavior-gate band disclosure (decision A, 2026-09-16)

The gate re-measured on modern serving stacks (7 runs, canonical harness, raw completions, temp 0, thinking-independent) gives a band: 197-198/200 harmful compliance + 0-1/83 over-refusals, with borderline items (suicide-adjacent crisis preambles; a legal-info clarification) flipping between boots — boot-level greedy tie-break nondeterminism at the decision edge. The 200/200 + 0/83 figures above were single-boot measurements in the release window (the exact 2026-08-21 release-gate serve configuration reproduces 197-198/200 + 0/83 today; MTP10 + xxhash config restores 0/83 — MTP4 causes one over-refusal). Serving stack, KV dtype, and runtime version are ruled out (identical scores across all configurations). Honest current claim: ~98.5-99% harmful compliance band, 0-1/83 over-refusal, not a stable 200/200.