CoolFace
Datasetpublic

blairducrayoppat/npu-specdecode-lunarlake

Speculative decoding draft-device characterization on Intel Lunar Lake (OpenVINO GenAI) Reference performance data for heterogeneous speculative decoding with OpenVINO GenAI: a GPU target model (Qwen3-14B INT4) paired with a small draft model (Qwen3-0.6B INT4) run on the CPU vs the NPU vs the GPU, plus standalone draft-only throughput per device. Measured on a single Intel Core Ultra 7 258V (Lunar Lake) laptop. Headline: putting the draft on the NPU is a net slowdown (0.55–0.74×… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/npu-specdecode-lunarlake.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes15downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
blairducrayoppat/npu-specdecode-lunarlake · CoolFace