llmspeed/llm-speed-benchmarks
llm-speed: signed LLM inference-speed benchmarks Crowdsourced, cryptographically signed measurements of how fast large language models actually run: decode tokens per second, time to first token, and latency, across consumer GPUs, Apple Silicon, and hosted APIs, under one reproducible workload suite (suite-v1). Live data and bulk downloads: https://llm-speed.com/data Per-run permalink: https://llm-speed.com/r/<id> Methodology: https://llm-speed.com/methodology DOI:… See the full description on the dataset page: https://huggingface.co/datasets/llmspeed/llm-speed-benchmarks.
llm-speed: signed LLM inference-speed benchmarks
Crowdsourced, cryptographically signed measurements of how fast large language models actually run: decode tokens per second, time to first token, and latency, across consumer GPUs, Apple Silicon, and hosted APIs, under one reproducible workload suite (suite-v1).
- Live data and bulk downloads: https://llm-speed.com/data
- Per-run permalink: https://llm-speed.com/r/<id>
- Methodology: https://llm-speed.com/methodology
- DOI: https://doi.org/10.5281/zenodo.21254813 (Zenodo archive; resolves to the latest version)
- License: CC BY 4.0. Reuse freely, including commercially, with attribution and a link back.
What is in it
Each row is the headline result of one signed benchmark run.
Full per-workload detail (prefill tok/s, TTFT, quantization, percentiles) for any run is at https://llm-speed.com/r/<id> and via the live API at https://api.llm-speed.com/v1/results/<id>. Files: runs.csv, runs.jsonl, runs.json, plus meta.json (counts and generation timestamp).
How it is measured
Every run is produced by the open-source llm-speed CLI under the suite-v1 workload set, fingerprinted to real hardware, and cryptographically signed on upload (EdDSA/JWS) so a number cannot be edited after the fact. The values are measured, not modeled. If two rows disagree for the same model and hardware, both are real submissions from different setups; open each permalink to see the raw workload detail.
Citation
llm-speed. The llm-speed dataset: signed LLM inference-speed benchmarks.
Zenodo (2026). https://doi.org/10.5281/zenodo.21254813. CC BY 4.0.