CoolFace
Datasetpublic

witcheer/windows-rtx-4060ti-8gb-bench-2026-05

Local LLM Bench — RTX 4060 Ti 8GB Real practitioner benchmarks of open-source LLMs on consumer 8GB VRAM hardware. Hardware GPU: NVIDIA GeForce RTX 4060 Ti (8GB VRAM) CPU: AMD Ryzen 5 7600X (6 cores, AM5) RAM: 32GB DDR5-6000 CL36 Platform: Windows 11 Runtime: LM Studio (CUDA backend) Methodology All models loaded with: Quantization: Q4_K_M (GGUF) Context length: 16384 tokens GPU offload: maximum (full GPU residency where it fits) Temperature: 0.7… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-bench-2026-05.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes15downloads
Dataset Card

Local LLM Bench — RTX 4060 Ti 8GB

Real practitioner benchmarks of open-source LLMs on consumer 8GB VRAM hardware.

Hardware

  • —GPU: NVIDIA GeForce RTX 4060 Ti (8GB VRAM)
  • —CPU: AMD Ryzen 5 7600X (6 cores, AM5)
  • —RAM: 32GB DDR5-6000 CL36
  • —Platform: Windows 11
  • —Runtime: LM Studio (CUDA backend)

Methodology

All models loaded with:

  • —Quantization: Q4KM (GGUF)
  • —Context length: 16384 tokens
  • —GPU offload: maximum (full GPU residency where it fits)
  • —Temperature: 0.7
  • —Top-p: 0.9

Single benchmark prompt (prompts/moe-explainer-500w.md):

Explain how mixture-of-experts (MoE) inference works in modern LLMs. Cover: routing mechanism, expert activation per token, why MoE is more efficient than dense models of equal parameter count, and practical tradeoffs for local deployment on consumer hardware. Aim for ~500 words.

For each model, captured:

  • —Tokens/sec (decode speed, warm)
  • —Time-to-first-token (TTFT)
  • —Tokens generated total
  • —VRAM used (model + KV cache)
  • —Quality score (1-5, subjective)
  • —Length-target adherence (hit / overshoot / undershoot)
  • —Notes on output character

Format

JSONL, one row per benchmark run. See data/bench-2026-05-05.jsonl. Full model responses in responses/<model-slug>/2026-05-05.md.

Intended use

  • —Reproducibility: methodology and prompts documented
  • —Reference for 8GB VRAM hardware planners
  • —Community contribution: add your own rows via PR

Limitations

  • —Single prompt — quality scores reflect performance on one task type (technical explanation). Coding, multilingual, long-context not tested here.
  • —Subjective quality scores. See responses/ for raw outputs to judge yourself.
  • —Llama-3.3-8B response contained fabricated benchmark numbers — flagged in notes; use with caution for factual tasks.

Citation

If you use this data, cite as:

Witcheer (2026). Local LLM Bench — RTX 4060 Ti 8GB. https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-bench-2026-05

Contact

https://huggingface.co/witcheer