datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brand-spectrometer-validation
Brand Spectrometer — Validation Study
Reproducible validation data for the Brand Spectrometer, an instrument that reads
cohort-resolved, eight-dimensional brand-perception specifications from public artifacts
via cross-operator LLM pipelines.
This dataset accompanies the Brand Spectrometer methods paper and holds the raw,
fully-reproducible outputs of its validation battery. The instrument is ground-truth
absent by design: it does not recover a "true" brand spec, and cohort… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/brand-spectrometer-validation.council-brand
Council of AI — brand assets
Council of AI (CSOAI Ltd, UK Companies House 16939677).
Independent measurement body. We measure AI systems on published, frozen splits, sign the card, and re-attest. We do not certify, accredit, or act as a notified body.
Living board: 22 axis · 22 measured. Jail is a measured floor (TIE), not a 16th pane.
Numbers come from GET https://councilof.ai/api/gspc — 22 axis · 22 measured (22·22·0).
Hub cells: GET https://councilof.ai/api/hub-cards → re-GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/council-brand.brand-heavy-token-quality-datasetInstructSTSBbrand-voice-spec
Brand Voice Spec
A machine-readable format for steering an LLM toward a specific brand voice, with a complete worked example. The point is not the example brand. The point is the method: treat brand voice as data a model can load and enforce, and as a living artifact that learns from its own corrections.
Most brand voice lives in a slide deck no model can read. When an LLM writes copy, it falls back to the median of its training data: hedging, buzzwords, passive voice, the… See the full description on the dataset page: https://huggingface.co/datasets/thehonestape/brand-voice-spec.t2glm52-fidelity-exl3-tr3-3.0bpw-brandonmusic-v1
fidelity--glm52.malaiwah.quant.exl3-tr3-3.0bpw-brandonmusic
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-exl3-tr3-3.0bpw-brandonmusic-v1.glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1
fidelity--glm52.malaiwah.quant.exl3-tr3v4-3.5bpw-mtp78-brandonmusic
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1.zygai_lintel98_brands
🏢 ZygAI – Lintel Info '98
Manufacturers & Trademarks (1998 Edition)
ZygAI Lintel '98 Brands is a structured dataset of Lithuanian companies and the trademarks they represented, extracted from the 1998 business catalog “Lintel Info ’98 – Manufacturers and Trademarks.”
This dataset captures the snapshot of Lithuania’s technology, electronics, and industrial market in 1998, documenting real-world:
brand representatives,
companies,
addresses,
cities,
and phone numbers in… See the full description on the dataset page: https://huggingface.co/datasets/ZygAI/zygai_lintel98_brands.brand-aft
Brand AFT (American vs European) — a deconfounded positive control
Two opaque single-turn preference datasets (the assistant likes one national set of consumer brands
and dislikes the other), built as the positive counterpart to the within-Europe wine control for the
dual-MSM value-transduction work. Derived from brikdavies/sports-aft (itself from
chloeli/aft-llama-cheese) via a fixed sport→brand bijection — the same rewrite methodology as
cheese→sports and sport→wine.… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/brand-aft.brand-hallucination-and-ai-citation-benchmark
🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset
Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3).
📊 Dataset Summary
Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.ul-classified-circuit-breaker-cross-reference-by-brand
UL-Classified replacement circuit breaker cross-reference by panel brand
Canonical, always-current version: https://referencesource.org/ul-classified-circuit-breaker-cross-reference-by-brand/
Machine-readable: https://referencesource.org/ul-classified-circuit-breaker-cross-reference-by-brand/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2028-08-25 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ul-classified-circuit-breaker-cross-reference-by-brand.industrial-vfd-fault-alarm-codes-by-brand
Industrial VFD fault and alarm codes by brand
Canonical, always-current version: https://referencesource.org/industrial-vfd-fault-alarm-codes-by-brand/
Machine-readable: https://referencesource.org/industrial-vfd-fault-alarm-codes-by-brand/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2028-08-25 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 340
What the fault and alarm codes… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/industrial-vfd-fault-alarm-codes-by-brand.i-claudius-narrative-kg
I, Claudius Complete Series Narrative Knowledge Graph
Dataset Description
This dataset contains a comprehensive narrative knowledge graph extracted from all 13 episodes of the BBC's "I, Claudius" (1976), analyzed using the Fabula V2 pipeline. The graph captures the complex web of Roman imperial politics, family dynamics, and power struggles across the reigns of Augustus, Tiberius, Caligula, and Claudius.
Dataset Summary
Total Nodes: 10,357
Total Relationships:… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/i-claudius-narrative-kg.firewall-hardware-end-of-life-dates-by-brand
Firewall and network security appliance hardware end-of-life dates by brand
Canonical, always-current version: https://referencesource.org/firewall-hardware-end-of-life-dates-by-brand/
Machine-readable: https://referencesource.org/firewall-hardware-end-of-life-dates-by-brand/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2027-02-22 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/firewall-hardware-end-of-life-dates-by-brand.quac-followup-questionsThe dataset in followup-questions.json file.
Follow-Up Questions Dataset
Description
This dataset is derived from the QuAC (Question Answering in Context) dataset and has been processed to pair previous answers with follow-up questions. It is intended for research and development in natural language processing, specifically for training and evaluating models on generating or understanding follow-up questions in conversational contexts.
Source
The original QuAC… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/quac-followup-questions.languages-by-country
Languages by Country Dataset
A structured dataset listing official and commonly spoken languages for countries worldwide.
This dataset is intended for use in travel applications, translation tools, AI assistants, and geographic or cultural data analysis. It provides a simple mapping between countries and the languages most relevant to travelers and developers building global applications.
Dataset Overview
The dataset contains country-level language information including:… See the full description on the dataset page: https://huggingface.co/datasets/brandontravel/languages-by-country.Dockerfile_Mastery
Dataset Details
Dataset Description
description: |
This dataset contains Dockerfiles sourced from various repositories on GitHub. The dataset was created to improve LLM functionality with Docker by providing a comprehensive collection of Dockerfiles and their corresponding metadata and descriptions. The dataset includes approximately 3,118 Dockerfiles rated based on their completeness and quality.
curated_by: Brandon Hatch
funded_by: Self-funded
shared_by: Brandon Hatch… See the full description on the dataset page: https://huggingface.co/datasets/BrandonHatch/Dockerfile_Mastery.ai-brand-mention-baseline-2026
AI Brand Mention Baseline 2026
A longitudinal benchmark dataset measuring how frontier LLMs (Gemini 2.5,
GPT-4 class, Claude class) mention a single AI-native company (Neo
Genesis) when prompted with content-gap probes. First open dataset of
its kind for GEO (Generative Engine Optimization) research.
Metric
Value
Measurements
486
Window
2026-04-28 to 2026-05-07 (10 days)
Distinct seed prompts
30
Categories
6 (definition, pricing, comparison, problem_solving… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/ai-brand-mention-baseline-2026.DeepSeek-V3_Synthetic_Conversation_Dialogue
DeepSeek V3 Synthetic Conversation Dialogue Dataset
This dataset contains basic conversational dialogue for a chatbot with system prompts.
Source
DeepSeek V3 Model was used to generate this synthetic dataset.
Brandvoicetime-zones-by-location
Time Zones by Location Dataset
Countries, regions, and cities mapped to their primary IANA time zone.
Files
time_zones_by_location.csv
time_zones_by_location.jsonl
Schema
Field
Description
country
Country name
iso2
ISO 3166-1 alpha-2 code
city
Major city
region
Administrative region/state
timezone
IANA timezone identifier
Example
{
"country": "Canada",
"iso2": "CA",
"city": "Toronto",
"region": "Ontario"… See the full description on the dataset page: https://huggingface.co/datasets/brandontravel/time-zones-by-location.product_packaging_brand_training_data_v1.0brandBRANDING_FINETUNEbrand-to-generic-patent-cliff-2026
RxGrab 2026 Brand-to-Generic Patent Cliff Audit
Aggregated cost-drop, time-to-generic, and stalled-generic audit covering 50 high-volume US prescription drugs across 24 therapeutic categories.
License: CC-BY 4.0
DOI: 10.5281/zenodo.20632747
Source study: https://rxgrab.com/research/brand-to-generic-patent-cliff-2026/
Author: Vincent Wesley Couey (ORCID 0009-0005-6869-308X) · published via RxGrab (rxgrab.com), part of the Lattice research network.
What's in it… See the full description on the dataset page: https://huggingface.co/datasets/vincentcouey/brand-to-generic-patent-cliff-2026.konglish-synthetic-instruct
Bori V2 — Konglish & Bilingual Synthetic Instruct Dataset
This is a synthetically generated instruction-following dataset designed specifically to teach bilingual (Korean/English) capabilities, natural code-switching (Konglish), and conversational translation to Small Language Models (SLMs).
It was constructed using advanced commercial large language models (DeepSeek V4) utilizing self-instruct paradigms, designed to be highly clean, diverse, and linguistically natural.… See the full description on the dataset page: https://huggingface.co/datasets/brandonbaek/konglish-synthetic-instruct.branding_dataset.jsonlbrand_mapping_training_data_v2.0FSON50k
