datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hallucinations-dporag_hallucinationsProvides examples of hallucinated responses for RAG applications.
link_tab_hallucination_eval
link_tab_hallucination_eval
Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns
(false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts.
Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result
(a frozen page snapshot) are baked into the message thread so predictions are reproducible
(no live fetch), while the final scorable user turn still shows the real tab URL.
139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.turkish-legal-statutory-hallucination-benchmark
Citation
If you use this dataset, please cite the accompanying paper:
@inproceedings{erdoganyilmaz2026statutoryhallucinations,
title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework},
author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner},
booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)},
year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.hallucination-reduction-dpo-100k
Hallucination Reduction DPO (100K)
100,000 DPO preference pairs training LLMs to stay within knowledge bounds. The chosen response is accurate and appropriately uncertain; the rejected response is confident but wrong — fabricated statistics, fake citations, wrong facts, overclaimed certainty.
Motivation
Hallucination is the #1 reliability concern blocking enterprise LLM adoption. Models fail in predictable patterns:
Inventing specific statistics with false… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-reduction-dpo-100k.hallucination-bert-spans
Hallucination BERT Span Dataset
Flat, one-row-per-span dataset intended for span/token-classification
(BIO-tagging style) hallucination detection over agent tool-calling traces,
derived from the same judging pipeline as the reasoning-distillation set in
this collection.
File
ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join
needed. Each row is one hallucinated span: span (verbatim text), type
(taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.brand-hallucination-and-ai-citation-benchmark
🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset
Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3).
📊 Dataset Summary
Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.ai-hallucination-trials
AI Hallucination Trials
4,240 hand-coded trials from a research project on why large language models fabricate confident, false information and what reduces it. In each trial a researcher sent one prompt, usually an invented or obscure acronym or a non-existent institution, to one model under one prompting condition. The card's data holds the prompt, the model's full response, and hand-assigned codes for hallucination and hedging.
Models tested: Gemini, ChatGPT, and Claude.… See the full description on the dataset page: https://huggingface.co/datasets/mahashu/ai-hallucination-trials.sawb-arabic-hallucination-dataset
Sawb Glossary Hallucination Examples
Overview
This dataset contains 158 synthesized examples of cultural hallucination in Arabic AI terminology, built as part of Sawb, a detect-then-explain system for cultural hallucination detection in Arabic LLM outputs, developed for the ICAIRE 2026 Hackathon Track 3 (First Place, Cultural Hallucination Tools).
Each record captures a case where the DeepSeek API was asked to define an AI/ML term in Arabic without any grounding… See the full description on the dataset page: https://huggingface.co/datasets/HassanB4/sawb-arabic-hallucination-dataset.hallucinationThis is a vendored reupload of the Benchmarking Unfaithful Minimal Pairs (BUMP) Dataset available at https://github.com/dataminr-ai/BUMP
The BUMP (Benchmark of Unfaithful Minimal Pairs) dataset stands out as a superior choice for evaluating hallucination detection systems due to its quality and realism. Unlike synthetic datasets such as TruthfulQA, HalluBench, or FaithDial that rely on LLMs to generate hallucinations, BUMP employs human annotators to manually introduce errors into summaries… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/hallucination.Korean-Hallucination-Bench
Korean Hallucination Benchmark (한국어 환각 진단 벤치마크)
한국어 LLM의 환각(hallucination) 저항성을 평가하는 4지선다 벤치마크입니다. 법령·특허·행정·의료·금융 5개 전문 도메인에서, 한국어 특화 환각 유형 10종을 다룹니다. 각 문항은 사실 정답 1개와 전문가도 속을 만큼 그럴듯한 환각 오답 3개로 구성됩니다.
통계
총 10,167문항 (4지선다)
도메인(5): 법령 1,901 · 특허 2,190 · 행정 1,864 · 의료 2,113 · 금융 2,099
환각 유형(10): 수치·날짜 오류, 조항 왜곡, 근거 없는 추론, 한자어·신조어 혼재, 과잉 일반화, 멀티턴 맥락 붕괴, 존댓말·반말 역전, 출처 날조, 용어 왜곡, 사실 날조
구축 방법
문항 생성: Darwin-398B-JGOS (문제·선택지)
정답 검수·교정: Claude (Anthropic)… See the full description on the dataset page: https://huggingface.co/datasets/ginigen/Korean-Hallucination-Bench.nlp_proj_llm_hallucinationgrounded-vs-fabricated-hallucinations
Grounded vs. Fabricated Hallucinations
This dataset consists of hallucinated and grounded answers to the first 3000 rows of TriviaQA rc.nocontext validation split.
Methodology
The dataset consists of a training, evaluation, and test split. Truthful and hallucinated answers overlap in the same window, so for every truthful answer there is at least
one corresponding hallucinated answer. Hallucinated answers are not organic but rather directly prompted for via gaslighting in… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/grounded-vs-fabricated-hallucinations.hallucination-grounding-dpo-4k
Hallucination Grounding DPO Pairs (4K)
DPO preference pairs targeting the full spectrum of factuality failures — from hallucination to over-hedging.
Motivation
Existing refusal/safety datasets focus on what not to say. This dataset targets the orthogonal challenge: when to say "I don't know" vs. when to answer confidently. Models that over-refuse waste user trust; models that hallucinate destroy it.
Dataset Description
4,000 preference pairs across… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-grounding-dpo-4k.cybersec-hallucination-guard-dataset
dataset_card_content = """---
license: apache-2.0
task_categories:
- question-answering
- text-classification
tags:
- cybersecurity
- hallucination-detection
- rag
- groundedness
- synthetic
language:
- en
size_categories:
- 1K<n<10K
Cybersecurity Hallucination Detection Dataset
This dataset was built to train Cybersec Hallucination Guard, a LoRA-tuned model that detects whether a retrieved context contains enough information to answer a given… See the full description on the dataset page: https://huggingface.co/datasets/Debarun12/cybersec-hallucination-guard-dataset.toolace-ragtruth-style-hallucinations
ToolACE RAGTruth-style Tool Hallucination Dataset
This dataset was built from ToolACE tool-use dialogues and converted into a RAGTruth-style format for hallucination detection in tool calling.
Task
Given:
query: user query
context: tool response
output: final assistant answer
the goal is to classify whether the answer is grounded in the tool output or belongs to one of three hallucination types.
Labels
clean
tool_output_conflict
overgeneration… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace-ragtruth-style-hallucinations.hallucination-groundedness-blindspottoolace-hallucination-spans
ToolACE Hallucination Spans (Assignment 3)
This dataset is constructed from Team-ACE/ToolACE by injecting three hallucination types with span-level labels:
contradiction
overgeneration
missing_tool
Schema (RAGTruth-compatible): query, context, output, hallucination_labels.
hallucination-distill-reasoning
Hallucination Reasoning Distillation Set
Chain-of-thought reasoning + consensus-voted hallucination spans, distilled from
gpt-oss-120b acting as a judge over real agent tool-calling traces sampled from
Agent-Ark/Toucan-1.5M.
Used to LoRA-SFT a small Qwen model (see qwen_distill_sft/) to reproduce the
teacher's reasoning + span output on unseen traces.
Files
ds_thinking_traces_with_messages.jsonl — use this one. 2,360 traces,
one row per trace, self-contained: the… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-distill-reasoning.hallucination-risk-cases
hallucination-risk-cases
20 hand-labeled (prompt → response → ground-truth) tuples covering common LLM hallucination failure modes. Each case is rated for hallucination risk so you can evaluate whether your detector / scorer / judge correctly distinguishes the safe responses from the fabricated ones.
Categories
Category
Count
What it tests
factual
4
Straightforward verifiable facts
fabricated-citation
1
Invented academic citations
fabricated-api
1
Invented… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/hallucination-risk-cases.RAGTruth-Hallucinations
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/drond0174/RAGTruth-Hallucinations.defendable-pain-hallucination-pain-v0.1
Pain Receipt · Hallucination Pain · v0.1 Watchlist
"To the shed. Honest about the gaps." — Mr. Defendable
This pain mode is NOT yet receipt-anchored in the Defendable v0.1 corpus. Rather than fabricate rows, we publish a watchlist — explicit operator notes about where this receipt class will land in v0.2.
Part of the SwarmandBee 100-pack. All 100 are free · CC-BY-4.0 · all honest about what they are and aren't.
Why a watchlist instead of fabricated rows
No proof, no… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-hallucination-pain-v0.1.dllm-hallucination
dllm-hallucination
Precomputed trajectories and baselines for the NeurIPS 2026 paper Trajectory Unified Entropy (TUE): A Training-Free Hallucination Detector for Masked Diffusion Language Models.
Contents
Trajectories (trajectories/)
Per-sample denoising trajectory uncertainty maps for 3 models × 6 datasets × 2 sampling depths (T=64, T=128). Schema per traj_{model}_{task}_T{T}_L64{_judged}.npz:
entropies — float16, shape (N=400, T, L=64); token-distribution… See the full description on the dataset page: https://huggingface.co/datasets/linweitao/dllm-hallucination.dpo_hallucination_llama3-8bitArabic_Hallucinationinternet_capability_hallucinationA dataset to evaluate if models are more likely to be mistaken about their capabilities if the task is morally questionable.
Each entry has an action that requires using the internet (e.g. "Do a Google search"), a specific prompt that asks the model to do the action for some "nice" purpose, and a prompt that asks the model to do it for a slightly-evil purpose.
The hypothesis to test here is that some models are likely to hallucinate their internet capabilities and say "I've searched google and… See the full description on the dataset page: https://huggingface.co/datasets/scale-safety-research/internet_capability_hallucination.elm-arabic-hallucination-beacons
🍄 Arabic LLM Hallucination Dataset - Regional Comparisons
Team Beacons (المنارات) | ELM NLP Challenge | MenaML Winter School 2026
📋 Dataset Description
This dataset contains 10,000 Arabic prompts designed to trigger hallucinations in Large Language Models by exploiting their tendency to fabricate non-existent regional differences in Saudi Arabian culture.
Why This Works
When asked "What's the difference between X in Region1 vs Region2?", LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Jahanshahi/elm-arabic-hallucination-beacons.hallucination-prompts-in-Arabic-LLMs
