datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.speculative-decoding-bench-rtx4090
Speculative Decoding Benchmark — RTX 4090
TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.nanochat-rtx4070-sft-mixes
nanochat-rtx4070 SFT mixes
Eight SFT data mixes that were trained and evaluated on a single RTX 4070, and the results each one produced. Seven of them failed.
These are the actual independent variable behind the negative-results table in Bl4ckd09/nanochat-on-rtx4070. Every mix here was built deterministically, trained on the same frozen backbone with the same geometry and step count, and put through the same two-stage evaluation gate. Publishing only the winner would make the… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/nanochat-rtx4070-sft-mixes.vast-rtx3090-market-6mo
Vast.ai RTX 3090 Spot Market, February-August 2026
Panel data from the vast.ai GPU rental marketplace, restricted to NVIDIA RTX 3090 offers. The public offer listing was polled every 10 minutes between 2026-02-13 and 2026-08-15. Each observation records price, hardware specifications, host reliability, and location. A derived lifecycle table gives the listing duration of every offer. Vast.ai does not publish historical listing data; this dataset was collected independently.… See the full description on the dataset page: https://huggingface.co/datasets/MarcusLammers/vast-rtx3090-market-6mo.windows-rtx-4060ti-8gb-moe-offload-bench-2026-05
RTX 4060 Ti 8GB — Multi-Model Benchmark (2026-05)
practitioner benchmarks on consumer hardware (8GB VRAM, 32GB RAM). 10 models tested, covering MoE expert offload, hybrid SSM architectures, dense models, MLA, dense partial GPU offload, and the 1B speed ceiling. all runs on the same physical rig, same methodology.
current leaderboard (decode tok/s at sweet spot)
model
active params
GGUF size
sweet spot tok/s
quality (6 tests)
architecture
Llama 3.2 1B
1.24B
771… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-moe-offload-bench-2026-05.rtx-5080-llm-power-efficiency
RTX 5080 LLM Power Efficiency: Watts and Joules per Token
Measured board power, tokens per joule, and electricity cost per million tokens for local LLM inference on a retail RTX 5080, with 1Hz telemetry.
What this measures
Board power logged with nvidia-smi at 1Hz while driving fixed-length generations on a retail RTX 5080, converted to tokens per joule and dollars per million tokens.
Efficiency
Model
tokens/joule
median board W
Llama 3.2… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-power-efficiency.rtx-5080-local-llm-coding-eval
Local LLM Coding Evaluation on an RTX 5080
Task-level coding eval results for local models on a single RTX 5080, published with the open task set used to produce them.
Task-level results for local coding models measured on a single retail RTX 5080.
Published together with the open task set (coding-eval-tasks-v1.json) so the evaluation can be re-run or extended rather than merely cited. An eval you cannot reproduce is an anecdote.
Source and license
Canonical page… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-local-llm-coding-eval.omni-neural-sca-rtx3090-telemetry
Omni-Sovereign Neural SCA Physical Telemetry & Benchmark Dataset (NVIDIA GeForce RTX 3090)
Overview
This repository hosts the official, 100% empirical, zero-mock hardware side-channel telemetry datasets and cryptanalytic evaluation artifacts produced by the Hierarchical TCN-Attention Neural SCA Engine on physical NVIDIA GeForce RTX 3090 (Ampere GA102 / SM_86) silicon.
Rigorous 3-Tier Scientific Peer-Review Framework
To ensure scientific integrity… See the full description on the dataset page: https://huggingface.co/datasets/bbkdevops/omni-neural-sca-rtx3090-telemetry.rtx-5080-llm-throughput
RTX 5080 Local LLM Throughput — Measured, Not Estimated
Measured decode and prefill tokens/second, load times, and VRAM residency for
local LLMs on a single retail NVIDIA GeForce RTX 5080 (16 GB): gpt-oss:20b
(MXFP4), Qwen 2.5 14B and 7B (Q4_K_M), and Llama 3.2 3B (Q4_K_M), at 4k and
16k context with ~500 and ~1,650-token prompts.
Method
Every figure is the median of 3 runs using ollama's native
eval_count/eval_duration counters — never wall-clock division.
Every… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-throughput.rtx-4060ti-8gb-turboquant-bench-2026-05
RTX 4060 Ti 8GB — turboquant KV cache benchmark (Qwen3.6-35B-A3B)
practitioner-tested benchmarks of turboquant KV cache types vs standard llama.cpp on an RTX 4060 Ti 8GB with 32GB DDR5-6000 RAM.
hardware
component
spec
GPU
NVIDIA RTX 4060 Ti, 8 GB VRAM
CPU
AMD Ryzen 5 7600X (6c/12t)
RAM
32 GB DDR5-6000 dual-channel
OS
Windows 11 + WSL2 Ubuntu 26.04
model
Qwen3.6-35B-A3B-UD-Q4_K_M (22.1 GB). hybrid SSM+attention architecture — 10/40 layers… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-4060ti-8gb-turboquant-bench-2026-05.windows-rtx-4060ti-8gb-bench-2026-05
Local LLM Bench — RTX 4060 Ti 8GB
Real practitioner benchmarks of open-source LLMs on consumer 8GB VRAM hardware.
Hardware
GPU: NVIDIA GeForce RTX 4060 Ti (8GB VRAM)
CPU: AMD Ryzen 5 7600X (6 cores, AM5)
RAM: 32GB DDR5-6000 CL36
Platform: Windows 11
Runtime: LM Studio (CUDA backend)
Methodology
All models loaded with:
Quantization: Q4_K_M (GGUF)
Context length: 16384 tokens
GPU offload: maximum (full GPU residency where it fits)
Temperature: 0.7
Top-p: 0.9… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-bench-2026-05.finance_quiz_ko
