CoolFace
Datasetpublic

Zeraix/imparo-benchmarks

Imparo inference benchmarks Published performance measurements for Imparo, a hardware- and workload-adaptive LLM inference engine. This dataset makes the project's published benchmark table available in a machine-readable form. It contains 48 aggregate results across four engines and 12 model/workload combinations, not 48 independent benchmark runs or a training corpus. Project · Pinned source table · Zeraix organization Apple M3 Pro — published 2026-09-07… See the full description on the dataset page: https://huggingface.co/datasets/Zeraix/imparo-benchmarks.

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
0likes43downloads
Dataset Card

Imparo inference benchmarks

Published performance measurements for Imparo, a hardware- and workload-adaptive LLM inference engine.

This dataset makes the project's published benchmark table available in a machine-readable form. It contains 48 aggregate results across four engines and 12 model/workload combinations, not 48 independent benchmark runs or a training corpus.

Project · Pinned source table · Zeraix organization

Apple M3 Pro — published 2026-09-07

  • —Models: Gemma 4 E4B (UD-Q4KXL) and LFM2.5-2.6B (Q8_0).
  • —Engines: Imparo, llama.cpp, oMLX, and rapid-mlx.
  • —Configuration reported by the source: f16 KV cache, 512-token prefill chunks, thinking disabled on every engine.
  • —Method: two interleaved rounds in the same session; the client sends the same prompt to each engine in turn. The table reports the median in tokens/second.
  • —Timing: prefill throughput is derived from time to first token; decode throughput is derived from the token stream. These are client-observed measurements, not engine-reported kernel throughput.

Higher tokens/second is better. Compare engines within the same row, not across different models or prompt lengths.

ModelPrompt tokensPhaseImparollama.cppoMLXrapid-mlx
Gemma 4 E4B449prefill932556590830
Gemma 4 E4B5651prefill1063570904960
Gemma 4 E4B16191prefill1006522888914
LFM2.5-2.6B455prefill1023933675857
LFM2.5-2.6B5963prefill10609809641007
LFM2.5-2.6B17123prefill972900921957
Gemma 4 E4B449decode46.640.543.942.8
Gemma 4 E4B5651decode44.238.541.840.5
Gemma 4 E4B16191decode4034.537.936.7
LFM2.5-2.6B455decode46.643.946.844.6
LFM2.5-2.6B5963decode45.142.444.642.6
LFM2.5-2.6B17123decode42.74040.839.6

Scope and limitations

These values are transcribed from the public source table, not a new benchmark run or an independent reproduction. They describe the reported hardware, models, and configurations only.

The source table does not provide full per-round traces, exact prompt text, all tested engine commit IDs, RAM capacity, or OS version. The linked README commit pins the documentation snapshot, not the versions of every engine tested. Two rounds do not establish statistical significance; small differences should not be treated as definitive rankings. In particular, LFM2.5-2.6B short-prompt decode is 46.6 tok/s for Imparo and 46.8 tok/s for oMLX, described as a tie in the source.

M4 Pro: the project's 2026-09-10 update reports validation reproducing the optimization benefits observed on M3 Pro. No complete four-engine M4 Pro numerical table is included here. M3 Pro measurements must not be relabeled as M4 Pro results.

Data format

The JSONL file has one record per model, prompt length, phase, and engine:

FieldMeaning
measurement_idStable identifier for this published aggregate record.
hardwareHardware label reported in the source.
modelModel name reported in the source.
quantizationQuantization label reported for the model.
prompt_tokensPrompt length reported in the source table.
phaseprefill or decode.
engineEngine being measured.
tokenspersecondPublished median throughput in tok/s.

The configuration name uses the publication date, not a claimed exact execution timestamp. Future hardware results should be added as separate configurations, preserving earlier measurements and their provenance.

Contribute a reproducible result

Open a discussion or an Imparo issue. Include hardware and RAM, OS, model revision and quantization, engine commits, launch commands, prompts and output lengths, cache/warm-up settings, repeated runs, and correctness checks.

Attribution and license

Maintained by Zeraix. The source table comes from the Apache-2.0-licensed Imparo repository. This repository's documentation and benchmark-table transcription use Apache-2.0. No model weights, private logs, or user conversations are included. Upstream models and engines retain their own licenses.