datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.hallucinations-dpoWhisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a benchmark. Every evaluation config is test — do not fine-tune on it.
lexicon_synth is the exception: synthetic training material with its own train/test
split, and not one of the eight benchmark arms.
To build training data, exclude the items in
benchmark/exclusions.json
(546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The
benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.wiki_bio_gpt3_hallucination
Dataset Card for WikiBio GPT-3 Hallucination Dataset
GitHub repository: https://github.com/potsawee/selfcheckgpt
Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Dataset Summary
We generate Wikipedia-like passages using GPT-3 (text-davinci-003) using the prompt: This is a Wikipedia passage about {concept} where concept represents an individual from the WikiBio dataset.
We split the generated passages into… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/wiki_bio_gpt3_hallucination.rag_hallucinationsProvides examples of hallucinated responses for RAG applications.
ragtruth-translated-hallucinations
RAGTruth Translated Hallucinations
Multilingual machine translation of
RAGTruth into 31 European languages,
preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of
LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the
exact spans that are hallucinated (unsupported by, or contradicting, the provided
context). Here both the RAG prompt and the response are translated into each target
language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.lettucedetect-code-hallucination
LettuceDetect Grounded Hallucination Dataset
Token-level hallucination annotations on LLM responses grounded in structured
context across five sources — source code, developer-tool output, academic
papers, GitHub READMEs, and Wikipedia. Part of the LettuceDetect data
collection.
Every sample pairs a grounded context with an LLM answer that is either correct
or contains a minimally perturbed, character-span-annotated hallucination. All
spans use one unified taxonomy, so the… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/lettucedetect-code-hallucination.multi-wiki-qa-synthetic-hallucinationswhisper-hallucinations
Whisper Hallucinations on Noise
Dataset Summary
This dataset lists common hallucinations from OpenAI Whisper when the input has no speech.
We build it from a noise-only corpus.
We run Whisper on noise clips.
We collect any non-empty text that Whisper outputs.
We deduplicate phrases and count how often they occur.
Use it to test, detect, and reduce non-speech hallucinations.
Motivation
ASR models often output text on silence or noise.
These false hits harm UX… See the full description on the dataset page: https://huggingface.co/datasets/sachaarbonel/whisper-hallucinations.hallucination-guard-cachePhantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/seyled/Phantom_Hallucination_Detection.link_tab_hallucination_eval
link_tab_hallucination_eval
Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns
(false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts.
Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result
(a frozen page snapshot) are baked into the message thread so predictions are reproducible
(no live fetch), while the final scorable user turn still shows the real tab URL.
139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.hallucination-autopsy-benchmark
Hallucination Autopsy Benchmark
A unified, standardized benchmark for cross-model, cross-parameter analysis of LLM hallucination phenomena.
Overview
This dataset merges multiple hallucination detection benchmarks into a single standardized schema, enabling systematic etiological analysis of why and how different LLM architectures hallucinate under specific configurations.
Version: 3.0.0Total Records: 69,002Base Records: 69,002Augmented Records: 0Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/hallucination-autopsy-benchmark.FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/kankshith123/FinQA-hallucination-detection.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.rag-hallucination-dataset-1000
Retrieval-Augmented Generation (RAG) Hallucination Dataset 1000
Retrieval-Augmented Generation (RAG) Hallucination Dataset 1000 is an English dataset designed to reduce the hallucination in RAG-optimized models, built by Neural Bridge AI, and released under Apache license 2.0.
Dataset Description
Dataset Summary
Hallucination in large language models (LLMs) refers to the generation of incorrect, nonsensical, or unrelated text that does not stem from an… See the full description on the dataset page: https://huggingface.co/datasets/neural-bridge/rag-hallucination-dataset-1000.zip-training-hallucination-data-qwen06b-thinking-train-with-valuestoolace-unified-hallucinations
ToolACE Unified Hallucination Dataset
This repository contains a unified ToolACE-derived dataset for tool-calling hallucination research.
Files
data/train-00000-of-00001.parquet: leakage-safe grouped training split;
data/test-00000-of-00001.parquet: leakage-safe grouped test split.
The split was rebuilt at the normalized dialogue_id level,the same ToolACE dialogue can't appear in different splits.
Schema
column
description
system
system… See the full description on the dataset page: https://huggingface.co/datasets/VirVen/toolace-unified-hallucinations.toolace-tool-calling-hallucination-ragtruth
ToolACE-derived Tool-Calling Hallucination Dataset
This dataset was created for the course assignment Hallucination Detection in Tool Calling.
It is synthetic by design: starting from ToolACE-style tool-calling dialogues, we automatically inject three required hallucination types:
tool_contradiction
overgeneration
missing_tool
Each example follows a RAGTruth-like format:
query: user query
context: tool output
output: final model answer
hallucination_labels: span-level… See the full description on the dataset page: https://huggingface.co/datasets/marrita/toolace-tool-calling-hallucination-ragtruth.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.turkish-legal-statutory-hallucination-benchmark
Citation
If you use this dataset, please cite the accompanying paper:
@inproceedings{erdoganyilmaz2026statutoryhallucinations,
title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework},
author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner},
booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)},
year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.Taming-Hallucinations
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
CVPR 2026 Findings
Project Page | Paper | Code
Dataset Summary
This repository hosts DualityVidQA, the large-scale paired video–QA dataset introduced in
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation.
Taming Hallucinations introduces DualityForge, a controllable diffusion-based framework that turns
real videos into… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/Taming-Hallucinations.lettucedetect-prose-hallucination
LettuceDetect Prose Hallucination Dataset
Token-level hallucination annotations on LLM answers grounded in prose
context, drawn from two public RAG hallucination resources and mapped into one
unified taxonomy. This is the prose counterpart to the structured-context
(code, tool output, documents)
collection — together they let a single detector be trained across modalities.
Two sources sit side by side, distinguished by the dataset field:
dataset
Spans
Source
psiloqa… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/lettucedetect-prose-hallucination.Benchmark_Hallucinations_Datafinqa-data-processed-hallucination
FinQA Dataset with Hallucination Examples
The generated Weights & Biases Weave traces from this dataset generation process are publically available here.
Dataset Description
Dataset Summary
This dataset extends the original FinQA dataset by adding synthetic hallucinated examples for evaluating model truthfulness. Each original example is paired with a modified version that contains subtle hallucinations while maintaining natural language flow.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/wandb/finqa-data-processed-hallucination.curatorkit-testrun-Hallucination
curatorkit-testrun-Hallucination
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-28 10:47 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Hallucination", "alpaca")
hallucination-heads-longfact-augmentedlegal_hallucinations_paper_datalegal_rag_hallucinations
Dataset Card for Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools
This data release contains the queries and raw model outputs we analyze in Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies (2024, forthcoming).
Consistent with emerging understanding of AI
benchmarking and leaderboards, we reserve a random sample of 50% of the dataset to… See the full description on the dataset page: https://huggingface.co/datasets/reglab/legal_rag_hallucinations.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.
