CoolFace
9 results

rtx-5080

iBlessi /rtx-5080-llm-power-efficiency RTX 5080 LLM Power Efficiency: Watts and Joules per Token Measured board power, tokens per joule, and electricity cost per million tokens for local LLM inference on a retail RTX 5080, with 1Hz telemetry. What this measures Board power logged with nvidia-smi at 1Hz while driving fixed-length generations on a retail RTX 5080, converted to tokens per joule and dollars per million tokens. Efficiency Model tokens/joule median board W Llama 3.2… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-power-efficiency.tabularn<1K0 likes58 downloads28d agoHugging FaceiBlessi /rtx-5080-local-llm-coding-eval Local LLM Coding Evaluation on an RTX 5080 Task-level coding eval results for local models on a single RTX 5080, published with the open task set used to produce them. Task-level results for local coding models measured on a single retail RTX 5080. Published together with the open task set (coding-eval-tasks-v1.json) so the evaluation can be re-run or extended rather than merely cited. An eval you cannot reproduce is an anecdote. Source and license Canonical page… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-local-llm-coding-eval.textn<1K0 likes37 downloads28d agoHugging FaceiBlessi /rtx-5080-llm-throughput RTX 5080 Local LLM Throughput — Measured, Not Estimated Measured decode and prefill tokens/second, load times, and VRAM residency for local LLMs on a single retail NVIDIA GeForce RTX 5080 (16 GB): gpt-oss:20b (MXFP4), Qwen 2.5 14B and 7B (Q4_K_M), and Llama 3.2 3B (Q4_K_M), at 4k and 16k context with ~500 and ~1,650-token prompts. Method Every figure is the median of 3 runs using ollama's native eval_count/eval_duration counters — never wall-clock division. Every… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-throughput.tabularn<1K0 likes19 downloads2mo agoHugging Face