CoolFace
Modelpublic

Danny-Dasilva/Qwen3-8B-antidoom-W4A16-GPTQ-g32-DSpark

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes60downloads
Model Card

Qwen3-8B — Antidoom + W4A16 GPTQ (group=32) · DSpark-verified

An [Antidoom](https://github.com/Liquid4All/antidoom) (FTPO anti-repetition) version of `Qwen/Qwen3-8B`, quantized post-training to GPTQ int4 (W4A16, group=32, symmetric) and verified as a [DSpark](https://github.com/deepseek-ai/DeepSpec) speculative-decoding target (draft: `deepseek-ai/dspark_qwen3_8b_block7`).

341 tok/s single-stream on one RTX 5090 (32 GB) with DSpark k=7 + CUDA graphs — at 6.0 GB of weights.

What was done

  1. 1.Antidoom FTPO pass: sampled completions at temperature 0.01 from the LiquidAI/antidoom-mix-v1.0 prompt mix (filtered to the highest doom-loop-yield sources: code + instruction-following), detected runaway repetition, extracted 156 FTPO preference pairs (105 after regularisation filtering), trained a LoRA (r=128, all layers + lmhead, maxseq 3072) with early stop at chosen_win ≥ 0.4, and merged it into the base weights.
  1. 1.Post-training GPTQ (llm-compressor): int4 W4A16 group=32 symmetric, 256 ultrachat calibration samples with chat template. Post-training quantization preserves DSpark draft acceptance — retraining-style (QAT) quantization destroys it (see gemma-4-12B-it-W4A16-GPTQ-g32-DSpark).

DSpark acceptance is preserved through both steps: ~25% (original bf16) → 24.8% (after antidoom) → 26.1% (after GPTQ). The FTPO patch is local and targeted; it does not break the draft head.

Measured speed — RTX 5090 (32 GB, Blackwell), single stream, greedy, 256-tok gens

configtok/sDSpark accept
this model + DSpark k=7 + CUDA graphs341 (peaks 454)26.1%
this model, native (no speculation)224—
antidoom merged bf16 + DSpark k=718824.8%

Usage (vLLM + DSpark)

python
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer

MODEL = "Danny-Dasilva/Qwen3-8B-antidoom-W4A16-GPTQ-g32-DSpark"
DRAFT = "deepseek-ai/dspark_qwen3_8b_block7"

llm = LLM(
    model=MODEL,
    max_model_len=8192,
    attention_backend="FLASHINFER",
    speculative_config={
        "method": "dspark",
        "model": DRAFT,
        "num_speculative_tokens": 7,   # draft block size; keep at 7
        "attention_backend": "TRITON_ATTN",
    },
)
tok = AutoTokenizer.from_pretrained(MODEL)
prompt = tok.apply_chat_template([{"role": "user", "content": "Hello!"}],
                                 add_generation_prompt=True, tokenize=False)
print(llm.generate([prompt], SamplingParams(temperature=0.0, max_tokens=256))[0].outputs[0].text)

Requires a vLLM build with DSpark support (merged to vLLM main 2026-07-02).

Antidoom details

  • —Pipeline: Liquid4All/antidoom at upstream defaults except: temperature=0.01, prompt sources filtered to code/instruction-following (3-8× higher loop yield than math/QA), pair budget 156 (generation cut at ~5h; the chosen-win early stop converges regardless), num_epochs=3, lora_r=128, max_seq_length=3072 (fits 32 GB — the FTPO trainer materialises full-vocab logits), with the built-in chosen_win ≥ 0.4 early stop.
  • —The base model is healthy — only ~2% of low-temperature completions doom-loop — so this is a targeted robustness patch against runaway repetition, not a rescue. FTPO trains only local single-token preferences at loop-start positions.

Provenance

  • —Qwen/Qwen3-8B (bf16) → antidoom FTPO LoRA merge → GPTQ W4A16 g32
  • —Built 2026-07-07 on a single RTX 5090. Apache-2.0, same as the base model.