datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Witch_Bowl
Dataset Card for The Cauldron
Dataset description
The Cauldron is part of the Idefics2 release.
It is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
Load the dataset
To load the dataset, install the library datasets with pip install datasets. Then,
from datasets import load_dataset
ds = load_dataset("HuggingFaceM4/the_cauldron", "ai2d")
to download and load the… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/Witch_Bowl.strike-witches-501strtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.Void-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.agentic-score-leaderboard
🛠️ Agentic Score Leaderboard — one RTX 5090
How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
Leaderboard
#
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.local-agentic-coding-bench-8gb-vram-2026-05
agentic coding benchmark: local LLMs on 8GB VRAM
can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results.
hardware
GPU: NVIDIA RTX 4060 Ti 8GB
CPU: Intel i7-14700F
RAM: 32 GB DDR5
OS: Windows 11 + WSL2 (Ubuntu)
inference: llama-server (turboquant fork of llama.cpp)
what was tested
two agent frameworks:
Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.windows-rtx-4060ti-8gb-moe-offload-bench-2026-05
RTX 4060 Ti 8GB — Multi-Model Benchmark (2026-05)
practitioner benchmarks on consumer hardware (8GB VRAM, 32GB RAM). 10 models tested, covering MoE expert offload, hybrid SSM architectures, dense models, MLA, dense partial GPU offload, and the 1B speed ceiling. all runs on the same physical rig, same methodology.
current leaderboard (decode tok/s at sweet spot)
model
active params
GGUF size
sweet spot tok/s
quality (6 tests)
architecture
Llama 3.2 1B
1.24B
771… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-moe-offload-bench-2026-05.Viking_Witch_flirty_and_erotic_behavior
Dataset Card for Viking Witch Flirty and Erotic Behavior (NSFW)
Disclaimer
Warning: Adult Content
This dataset contains explicit adult material, including themes of sensuality, eroticism, and mature content inspired by Norse mythology and role-playing scenarios. It is intended solely for individuals who are 18 years of age or older and who consent to and approve of Not Safe For Work (NSFW) erotic adult content.
If you are under 18, find such material offensive, or are not… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Viking_Witch_flirty_and_erotic_behavior.Witcher-GRPO-promptssovereign-asr-bench
Sovereign ASR Bench — RTX 5090
Local, self-hosted automatic speech recognition benchmarks on one RTX 5090 32GB.
Part of the WITCHEER local-AI rig. Methodology that matters: load-once measurement (so
RTFx times transcription, not model load), one shared text normalizer applied to every
model output and reference, and micro-averaged WER (total errors / total reference
words — the LibriSpeech standard). The board lives as data in board.csv (shown in the viewer).
Board —… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/sovereign-asr-bench.WitChathermes-pairing-bench
Hermes Pairing — Agentic Benchmark for Local LLMs (Phase A + B)
How well does a local LLM drive an agent? This dataset holds results for pairing local models with
Hermes Agent (NousResearch) — a CodeAct agent: the model
acts by writing Python (execute_code) that orchestrates tools, not by emitting JSON function calls.
Generated with llm-bench-rig on an NVIDIA RTX 5090 (32GB),
llama.cpp / GGUF, under Hermes's real ~3.5K-token system prompt.
Phase A (synthetic). A reproducible… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/hermes-pairing-bench.ridiculous_math_questions
Dataset Card for Ridiculous Math Questions
A Set of ridiculous math questions that you won't find a teacher to write!
Dataset Details
Dataset Description
This dataset is a list of math questions generated by large language a model.
Which model is used depends on the version:
v0.05 was written by a 20B model, specifically DaringMaid-20B-V1.1-6bpw-exl2.
Curated by: KaraKaraWitch
Funded by [optional]: N/A
Shared by [optional]: KaraKaraWitch
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/WitchesSocialStream/ridiculous_math_questions.witcherbot-datawitchspeech
WitchSpeech: Russian voice lines Witcher 3 TTS dataset
This is a repack of the dataset
so-vits-svc-4.0-ru-The_Witcher_3_Wild_Hunt for TTS purposes.Unlike the original dataset, this dataset also contains transcriptions for voice lines.Transcriptions include stresses for words, even for words with just a single vowel.Additionally, there is a metadata_source.csv file, that contains voice lines text “as-is”.
Dataset info (see details in stats.txt):
Sample rate: 48 000Total time:… See the full description on the dataset page: https://huggingface.co/datasets/korovsky/witchspeech.Smol-Witcher-pretrainingLKF-unlearning_Salem_Witch_TrialsHowItsMade
Dataset Card for How It's Made
This tiny dataset contains parsed subtitles from the Canadian documentary series: "How It's Made".
Dataset Details
Uses
This dataset is intended to be used in Large language models for grounding questions asking on "How X item is made?"
Direct Use
N/A. Dataset released As-Is.
Out-of-Scope Use
The auther thinks that this can be used to generate inaccurate descriptions.
Dataset Structure
Refer to the… See the full description on the dataset page: https://huggingface.co/datasets/WitchesSocialStream/HowItsMade.tokenized_dataset_bart_fblarge
Dataset Card for "tokenized_dataset_bart_fblarge"
More Information needed
Witcher-fandom-instruct-datasetWitcher-multilingual-convosrtx-4060ti-8gb-turboquant-bench-2026-05
RTX 4060 Ti 8GB — turboquant KV cache benchmark (Qwen3.6-35B-A3B)
practitioner-tested benchmarks of turboquant KV cache types vs standard llama.cpp on an RTX 4060 Ti 8GB with 32GB DDR5-6000 RAM.
hardware
component
spec
GPU
NVIDIA RTX 4060 Ti, 8 GB VRAM
CPU
AMD Ryzen 5 7600X (6c/12t)
RAM
32 GB DDR5-6000 dual-channel
OS
Windows 11 + WSL2 Ubuntu 26.04
model
Qwen3.6-35B-A3B-UD-Q4_K_M (22.1 GB). hybrid SSM+attention architecture — 10/40 layers… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-4060ti-8gb-turboquant-bench-2026-05.Witcher-synth-multi-round-instructhybrid_data_fin
Dataset Card for "hybrid_data_fin"
More Information needed
ada_002_embeddings
Dataset Card for "ada_002_embeddings"
More Information needed
tokenized_dataset_bart
Dataset Card for "tokenized_dataset_bart"
More Information needed
LKF-unlearning_Salem_Witch_trials_rephrasings_finalwitcher3-dataset-ptbrtokenized_T5_base
Dataset Card for "tokenized_T5_base"
More Information needed
Witcher-synth-instruct-dataset
