CoolFace
Modelpublic

AngelSlim/Gpt-oss-20b-dflare

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes53downloads
Model Card

DFlare Draft Model for GPT-OSS-20B

This is the official DFlare draft model checkpoint for `openai/gpt-oss-20b`, released alongside the paper:

DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding

DFlare is a block-diffusion speculative decoding framework that accelerates large language model inference by predicting an entire block of tokens in one shot for the target model to verify in parallel. It removes the narrow conditioning bottleneck of the prior state-of-the-art DFlash through a lightweight layer-wise fusion mechanism: each draft layer attends to its own learnable combination of a broad set of target layers at negligible overhead, simultaneously injecting richer target knowledge and giving every draft layer a distinct input. Combined with training-data scaling, this enhanced per-layer expressiveness allows the draft model to scale to deeper architectures with consistent gains, achieving 3.91ร— end-to-end speedup on GPT-OSS-20B without compromising output quality.

๐Ÿ“– Documentation & code: https://angelslim.readthedocs.io/zh-cn/latest/features/speculative_decoding/dflare.html

๐Ÿ“ฆ Repo: Tencent/AngelSlim


Model Details

Target model`openai/gpt-oss-20b`
Draft architectureDFlare (8 layers, hiddensize=2880, attentionheads=64, GQA kv_heads=8)
Parameters~767 M
Block size8
Target layers used for fusion[1, 4, 7, 12, 15, 18, 21] (out of 24)
Precisionbfloat16
RoPEyarn scaling (factor=32.0, original_max_position_embeddings=4096)
Vocab size201,088

The draft predicts a block of block_size tokens in parallel, conditioned on (i) target hidden states extracted from the listed target layers and (ii) noise embeddings of the previous block. The target model verifies the block in a single forward pass and accepts the longest matching prefix.


How to Use

This checkpoint is loaded with AngelSlim's QwenDFlareDraftModel class (it is a Qwen3-style draft architecture; the model_type: qwen3 field in config.json is correct).

1. Install AngelSlim

bash
git clone https://github.com/Tencent/AngelSlim.git
cd AngelSlim
pip install -e .

2. Run end-to-end speculative decoding benchmark

The repo ships a self-contained benchmark entry that supports both DFlash and DFlare drafts via --draft-arch:

bash
# Single-GPU
python tools/dflash_benchmark.py \
    --model-name-or-path openai/gpt-oss-20b \
    --draft-name-or-path dflare/gpt-oss-20b-dflare \
    --draft-arch dflare \
    --dataset gsm8k \
    --max-samples 128 \
    --max-new-tokens 2048 \
    --temperature 0.0
bash
# 8-GPU (workload sharded across ranks, results gathered to rank 0)
torchrun --nproc_per_node=8 --master_port=29600 \
    tools/dflash_benchmark.py \
    --model-name-or-path openai/gpt-oss-20b \
    --draft-name-or-path dflare/gpt-oss-20b-dflare \
    --draft-arch dflare \
    --dataset gsm8k \
    --max-samples 128 \
    --max-new-tokens 2048 \
    --temperature 0.0

The script reports:

  • โ€”Decoding speedup vs. single-token autoregressive decoding
  • โ€”Average acceptance length per block
  • โ€”Per-block acceptance-length histogram
โš ๏ธ Do not pass --block-size โ€” the benchmark reads block_size=8 from this checkpoint's config.json and overriding it will break the train/test alignment.

Supported datasets out of the box: gsm8k, math500, aime24, aime25, alpaca, mt-bench, humaneval, mbpp, lbpp, swe-bench, livecodebench.

3. Load the checkpoint manually

python
import torch
from angelslim.compressor.speculative.train.models.draft.qwen_dflare import (
    QwenDFlareDraftModel,
)

draft = QwenDFlareDraftModel.from_pretrained(
    "dflare/gpt-oss-20b-dflare",
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
).cuda().eval()

print(draft.target_layer_ids)   # [1, 4, 7, 12, 15, 18, 21]
print(draft.block_size)          # 8
print(draft.mask_token_id)       # 200019

Performance

On six benchmarks spanning mathematical reasoning, code generation, and conversation, DFlare on GPT-OSS-20B delivers 3.91ร— average wall-clock speedup over single-token autoregressive decoding โ€” improving over DFlash by roughly 5%, with no degradation in output quality (the target model verifies every block, so the final distribution is identical to greedy decoding).

For full per-task results, ablations, and acceptance-length distributions, see the official documentation.


Citation

bibtex
@article{DFlare2026,
  title={DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding},
  author={Jiebin Zhang and Zhenghan Yu and Song Liu and Eugene J. Yu and Zheng Li and Dawei Zhu and Jiangshan Duo and Weimin Xiong and Yifan Song and Guanghua Yu and Jianchen Zhu and Sujian Li},
  journal={arXiv preprint arXiv},
  year={2026}
}

License

This checkpoint is released under the Apache 2.0 license, following the AngelSlim project. The target model openai/gpt-oss-20b retains its own license; consult the target model card before deployment.