CoolFace
Datasetpublic

blairducrayoppat/npu-specdecode-lunarlake

Speculative decoding draft-device characterization on Intel Lunar Lake (OpenVINO GenAI) Reference performance data for heterogeneous speculative decoding with OpenVINO GenAI: a GPU target model (Qwen3-14B INT4) paired with a small draft model (Qwen3-0.6B INT4) run on the CPU vs the NPU vs the GPU, plus standalone draft-only throughput per device. Measured on a single Intel Core Ultra 7 258V (Lunar Lake) laptop. Headline: putting the draft on the NPU is a net slowdown (0.55–0.74×… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/npu-specdecode-lunarlake.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes14downloads

blairducrayoppat/npu-specdecode-lunarlake · main · files are served by the source, never re-hosted here