datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brand-spectrometer-validation
Brand Spectrometer — Validation Study
Reproducible validation data for the Brand Spectrometer, an instrument that reads
cohort-resolved, eight-dimensional brand-perception specifications from public artifacts
via cross-operator LLM pipelines.
This dataset accompanies the Brand Spectrometer methods paper and holds the raw,
fully-reproducible outputs of its validation battery. The instrument is ground-truth
absent by design: it does not recover a "true" brand spec, and cohort… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/brand-spectrometer-validation.glm52-fidelity-exl3-tr3-3.0bpw-brandonmusic-v1
fidelity--glm52.malaiwah.quant.exl3-tr3-3.0bpw-brandonmusic
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-exl3-tr3-3.0bpw-brandonmusic-v1.glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1
fidelity--glm52.malaiwah.quant.exl3-tr3v4-3.5bpw-mtp78-brandonmusic
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1.brand-hallucination-and-ai-citation-benchmark
🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset
Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3).
📊 Dataset Summary
Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.i-claudius-narrative-kg
I, Claudius Complete Series Narrative Knowledge Graph
Dataset Description
This dataset contains a comprehensive narrative knowledge graph extracted from all 13 episodes of the BBC's "I, Claudius" (1976), analyzed using the Fabula V2 pipeline. The graph captures the complex web of Roman imperial politics, family dynamics, and power struggles across the reigns of Augustus, Tiberius, Caligula, and Claudius.
Dataset Summary
Total Nodes: 10,357
Total Relationships:… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/i-claudius-narrative-kg.ai-brand-mention-baseline-2026
AI Brand Mention Baseline 2026
A longitudinal benchmark dataset measuring how frontier LLMs (Gemini 2.5,
GPT-4 class, Claude class) mention a single AI-native company (Neo
Genesis) when prompted with content-gap probes. First open dataset of
its kind for GEO (Generative Engine Optimization) research.
Metric
Value
Measurements
486
Window
2026-04-28 to 2026-05-07 (10 days)
Distinct seed prompts
30
Categories
6 (definition, pricing, comparison, problem_solving… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/ai-brand-mention-baseline-2026.
