Danny-Dasilva/Qwen3-8B-antidoom-W4A16-GPTQ-g32-DSpark
160
Qwen3-8B — Antidoom + W4A16 GPTQ (group=32) · DSpark-verified
An [Antidoom](https://github.com/Liquid4All/antidoom) (FTPO anti-repetition) version of `Qwen/Qwen3-8B`, quantized post-training to GPTQ int4 (W4A16, group=32, symmetric) and verified as a [DSpark](https://github.com/deepseek-ai/DeepSpec) speculative-decoding target (draft: `deepseek-ai/dspark_qwen3_8b_block7`).
341 tok/s single-stream on one RTX 5090 (32 GB) with DSpark k=7 + CUDA graphs — at 6.0 GB of weights.
What was done
- Antidoom FTPO pass: sampled completions at temperature 0.01 from the LiquidAI/antidoom-mix-v1.0 prompt mix (filtered to the highest doom-loop-yield sources: code + instruction-following), detected runaway repetition, extracted 156 FTPO preference pairs (105 after regularisation filtering), trained a LoRA (r=128, all layers + lmhead, maxseq 3072) with early stop at
chosen_win ≥ 0.4, and merged it into the base weights.
- Post-training GPTQ (llm-compressor): int4 W4A16 group=32 symmetric, 256 ultrachat calibration samples with chat template. Post-training quantization preserves DSpark draft acceptance — retraining-style (QAT) quantization destroys it (see gemma-4-12B-it-W4A16-GPTQ-g32-DSpark).
DSpark acceptance is preserved through both steps: ~25% (original bf16) → 24.8% (after antidoom) → 26.1% (after GPTQ). The FTPO patch is local and targeted; it does not break the draft head.
Measured speed — RTX 5090 (32 GB, Blackwell), single stream, greedy, 256-tok gens
Usage (vLLM + DSpark)
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
MODEL = "Danny-Dasilva/Qwen3-8B-antidoom-W4A16-GPTQ-g32-DSpark"
DRAFT = "deepseek-ai/dspark_qwen3_8b_block7"
llm = LLM(
model=MODEL,
max_model_len=8192,
attention_backend="FLASHINFER",
speculative_config={
"method": "dspark",
"model": DRAFT,
"num_speculative_tokens": 7, # draft block size; keep at 7
"attention_backend": "TRITON_ATTN",
},
)
tok = AutoTokenizer.from_pretrained(MODEL)
prompt = tok.apply_chat_template([{"role": "user", "content": "Hello!"}],
add_generation_prompt=True, tokenize=False)
print(llm.generate([prompt], SamplingParams(temperature=0.0, max_tokens=256))[0].outputs[0].text)Requires a vLLM build with DSpark support (merged to vLLM main 2026-07-02).
Antidoom details
- Pipeline: Liquid4All/antidoom at upstream defaults except:
temperature=0.01, prompt sources filtered to code/instruction-following (3-8× higher loop yield than math/QA), pair budget 156 (generation cut at ~5h; the chosen-win early stop converges regardless),num_epochs=3,lora_r=128,max_seq_length=3072(fits 32 GB — the FTPO trainer materialises full-vocab logits), with the built-inchosen_win ≥ 0.4early stop. - The base model is healthy — only ~2% of low-temperature completions doom-loop — so this is a targeted robustness patch against runaway repetition, not a rescue. FTPO trains only local single-token preferences at loop-start positions.
Provenance
Qwen/Qwen3-8B(bf16) → antidoom FTPO LoRA merge → GPTQ W4A16 g32- Built 2026-07-07 on a single RTX 5090. Apache-2.0, same as the base model.
