dflash2
Qwen3.8-27B-DFlash2-GGUFQwen3.8-27B-DFlash2Qwen3.8-27B-DFlash2Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUFQwen3.8-27B-DFlash2-GGUFQwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2GLM-5.3-Flash-DFlash2Qwen3.8-27B-DFlash2-W4A16
Datasets
All datasets matching “dflash2”qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark
Qwen3.8-27B (MLX 4-bit) + DFlash2 speculative decoding on M4 Pro — benchmark recipe
This is a benchmark recipe, not redistributed weights. It records the exact
hardware, software, and commands used to measure a 2.06x generation-throughput
speedup with DFlash speculative decoding, and how to rerun it.
Result
HumanEval, 20 samples, max 256 new tokens, temperature 0 (greedy), reasoning
off, block size 5, paired baseline and DFlash under identical settings. Other… See the full description on the dataset page: https://huggingface.co/datasets/hamiejuice/qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark.llama-cpp-dflash2-cuda-bin
