rtx-5080
Datasets
All datasets matching “rtx-5080”rtx-5080-llm-power-efficiency
RTX 5080 LLM Power Efficiency: Watts and Joules per Token
Measured board power, tokens per joule, and electricity cost per million tokens for local LLM inference on a retail RTX 5080, with 1Hz telemetry.
What this measures
Board power logged with nvidia-smi at 1Hz while driving fixed-length generations on a retail RTX 5080, converted to tokens per joule and dollars per million tokens.
Efficiency
Model
tokens/joule
median board W
Llama 3.2… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-power-efficiency.rtx-5080-local-llm-coding-eval
Local LLM Coding Evaluation on an RTX 5080
Task-level coding eval results for local models on a single RTX 5080, published with the open task set used to produce them.
Task-level results for local coding models measured on a single retail RTX 5080.
Published together with the open task set (coding-eval-tasks-v1.json) so the evaluation can be re-run or extended rather than merely cited. An eval you cannot reproduce is an anecdote.
Source and license
Canonical page… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-local-llm-coding-eval.rtx-5080-llm-throughput
RTX 5080 Local LLM Throughput — Measured, Not Estimated
Measured decode and prefill tokens/second, load times, and VRAM residency for
local LLMs on a single retail NVIDIA GeForce RTX 5080 (16 GB): gpt-oss:20b
(MXFP4), Qwen 2.5 14B and 7B (Q4_K_M), and Llama 3.2 3B (Q4_K_M), at 4k and
16k context with ~500 and ~1,650-token prompts.
Method
Every figure is the median of 3 runs using ollama's native
eval_count/eval_duration counters — never wall-clock division.
Every… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-throughput.
