CoolFace
Datasetpublic

hamiejuice/qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark

Qwen3.8-27B (MLX 4-bit) + DFlash2 speculative decoding on M4 Pro — benchmark recipe This is a benchmark recipe, not redistributed weights. It records the exact hardware, software, and commands used to measure a 2.06x generation-throughput speedup with DFlash speculative decoding, and how to rerun it. Result HumanEval, 20 samples, max 256 new tokens, temperature 0 (greedy), reasoning off, block size 5, paired baseline and DFlash under identical settings. Other… See the full description on the dataset page: https://huggingface.co/datasets/hamiejuice/qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark.

sourceHugging Faceupdated 28d agoView on Hugging Face
0likes60downloads
3 commits on main
b955aca28d ago

Update to 20-sample HumanEval benchmark

hamiejuice
ce979a128d ago

Publish paired MLX DFlash2 benchmark

hamiejuice
cfb742028d ago

initial commit

hamiejuice