datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.Void-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.agentic-score-leaderboard
🛠️ Agentic Score Leaderboard — one RTX 5090
How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
Leaderboard
#
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.local-agentic-coding-bench-8gb-vram-2026-05
agentic coding benchmark: local LLMs on 8GB VRAM
can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results.
hardware
GPU: NVIDIA RTX 4060 Ti 8GB
CPU: Intel i7-14700F
RAM: 32 GB DDR5
OS: Windows 11 + WSL2 (Ubuntu)
inference: llama-server (turboquant fork of llama.cpp)
what was tested
two agent frameworks:
Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.sovereign-asr-bench
Sovereign ASR Bench — RTX 5090
Local, self-hosted automatic speech recognition benchmarks on one RTX 5090 32GB.
Part of the WITCHEER local-AI rig. Methodology that matters: load-once measurement (so
RTFx times transcription, not model load), one shared text normalizer applied to every
model output and reference, and micro-averaged WER (total errors / total reference
words — the LibriSpeech standard). The board lives as data in board.csv (shown in the viewer).
Board —… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/sovereign-asr-bench.hermes-pairing-bench
Hermes Pairing — Agentic Benchmark for Local LLMs (Phase A + B)
How well does a local LLM drive an agent? This dataset holds results for pairing local models with
Hermes Agent (NousResearch) — a CodeAct agent: the model
acts by writing Python (execute_code) that orchestrates tools, not by emitting JSON function calls.
Generated with llm-bench-rig on an NVIDIA RTX 5090 (32GB),
llama.cpp / GGUF, under Hermes's real ~3.5K-token system prompt.
Phase A (synthetic). A reproducible… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/hermes-pairing-bench.LKF-unlearning_Salem_Witch_Trialsrtx-4060ti-8gb-turboquant-bench-2026-05
RTX 4060 Ti 8GB — turboquant KV cache benchmark (Qwen3.6-35B-A3B)
practitioner-tested benchmarks of turboquant KV cache types vs standard llama.cpp on an RTX 4060 Ti 8GB with 32GB DDR5-6000 RAM.
hardware
component
spec
GPU
NVIDIA RTX 4060 Ti, 8 GB VRAM
CPU
AMD Ryzen 5 7600X (6c/12t)
RAM
32 GB DDR5-6000 dual-channel
OS
Windows 11 + WSL2 Ubuntu 26.04
model
Qwen3.6-35B-A3B-UD-Q4_K_M (22.1 GB). hybrid SSM+attention architecture — 10/40 layers… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-4060ti-8gb-turboquant-bench-2026-05.LKF-unlearning_Salem_Witch_trials_rephrasings_final
