datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
resultsrequestsFinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.hallucinations-dpowiki_bio_gpt3_hallucination
Dataset Card for WikiBio GPT-3 Hallucination Dataset
GitHub repository: https://github.com/potsawee/selfcheckgpt
Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Dataset Summary
We generate Wikipedia-like passages using GPT-3 (text-davinci-003) using the prompt: This is a Wikipedia passage about {concept} where concept represents an individual from the WikiBio dataset.
We split the generated passages into… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/wiki_bio_gpt3_hallucination.Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a benchmark. Every evaluation config is test — do not fine-tune on it.
lexicon_synth is the exception: synthetic training material with its own train/test
split, and not one of the eight benchmark arms.
To build training data, exclude the items in
benchmark/exclusions.json
(546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The
benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.rag_hallucinationsProvides examples of hallucinated responses for RAG applications.
ragtruth-translated-hallucinations
RAGTruth Translated Hallucinations
Multilingual machine translation of
RAGTruth into 31 European languages,
preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of
LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the
exact spans that are hallucinated (unsupported by, or contradicting, the provided
context). Here both the RAG prompt and the response are translated into each target
language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.llama2-hallucination-hidden-states
🧠 LLaMA-2 Hidden-State Hallucination Dataset
Repository: ShoaibSSM/llama2-hallucination-hidden-states
Task: Hallucination Detection via Transformer Internal Representations
Base Model: LLaMA-2-7B
Primary Labels: LLM-Judge + Hybrid Semantic Grounding
📌 Overview
This dataset contains layer-wise hidden states extracted from LLaMA-2-7B during question answering on SQuAD v2, along with structured hallucination labels.
Unlike traditional hallucination datasets that operate… See the full description on the dataset page: https://huggingface.co/datasets/ShoaibSSM/llama2-hallucination-hidden-states.hallucination-guard-cachelegal_hallucinations_subset
Legal Hallucinations Subset
Dataset Description
This is a curated subset of the reglab/legal_hallucinations dataset, containing up to 1000 randomly sampled rows for each of 6 specific legal reasoning tasks (5444 rows total).
The original dataset was created for the paper: Dahl et al., "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models," Journal of Legal Analysis (2024, forthcoming). Preprint: arxiv:2401.01301
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legal_hallucinations_subset.link_tab_hallucination_eval
link_tab_hallucination_eval
Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns
(false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts.
Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result
(a frozen page snapshot) are baked into the message thread so predictions are reproducible
(no live fetch), while the final scorable user turn still shows the real tab URL.
139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.lettucedetect-code-hallucination
LettuceDetect Grounded Hallucination Dataset
Token-level hallucination annotations on LLM responses grounded in structured
context across five sources — source code, developer-tool output, academic
papers, GitHub READMEs, and Wikipedia. Part of the LettuceDetect data
collection.
Every sample pairs a grounded context with an LLM answer that is either correct
or contains a minimally perturbed, character-span-annotated hallucination. All
spans use one unified taxonomy, so the… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/lettucedetect-code-hallucination.multi-wiki-qa-synthetic-hallucinationswhisper-hallucinations
Whisper Hallucinations on Noise
Dataset Summary
This dataset lists common hallucinations from OpenAI Whisper when the input has no speech.
We build it from a noise-only corpus.
We run Whisper on noise clips.
We collect any non-empty text that Whisper outputs.
We deduplicate phrases and count how often they occur.
Use it to test, detect, and reduce non-speech hallucinations.
Motivation
ASR models often output text on silence or noise.
These false hits harm UX… See the full description on the dataset page: https://huggingface.co/datasets/sachaarbonel/whisper-hallucinations.hallucination-autopsy-benchmark
Hallucination Autopsy Benchmark
A unified, standardized benchmark for cross-model, cross-parameter analysis of LLM hallucination phenomena.
Overview
This dataset merges multiple hallucination detection benchmarks into a single standardized schema, enabling systematic etiological analysis of why and how different LLM architectures hallucinate under specific configurations.
Version: 3.0.0Total Records: 69,002Base Records: 69,002Augmented Records: 0Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/hallucination-autopsy-benchmark.Phantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/seyled/Phantom_Hallucination_Detection.LCAR-Hallucination-Benchmark
LCAR Hallucination Benchmark
LCAR Hallucination Benchmark is a manually reviewed speech benchmark for
studying acoustic-grounding failures in LLM-based ASR. It contains two
500-utterance suites: controlled speech synthesized with IndexTTS2 and speech
derived from openly released corpora. The benchmark covers translation or
transliteration, spoken or text-prompt instruction execution, unsupported
repetition, and catastrophic deletion.
The benchmark is a targeted stress set. It is… See the full description on the dataset page: https://huggingface.co/datasets/aguangguang/LCAR-Hallucination-Benchmark.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.video-hallucinationhallucination-probe-eval-results
Hallucination probe evaluation results
Resumable generation artifacts for Ours-Probe, TruthPrInt, and VIB-Probe.
The directory layout mirrors the source repository under outputs/.
MMHal and ObjectHal results use 256-token generation and three seeds: 0, 42, 1337.
Incomplete seed artifacts may be replaced by completed runs from another server.
rag-hallucination-dataset-1000
Retrieval-Augmented Generation (RAG) Hallucination Dataset 1000
Retrieval-Augmented Generation (RAG) Hallucination Dataset 1000 is an English dataset designed to reduce the hallucination in RAG-optimized models, built by Neural Bridge AI, and released under Apache license 2.0.
Dataset Description
Dataset Summary
Hallucination in large language models (LLMs) refers to the generation of incorrect, nonsensical, or unrelated text that does not stem from an… See the full description on the dataset page: https://huggingface.co/datasets/neural-bridge/rag-hallucination-dataset-1000.legal_hallucinations
Dataset Card for Legal Hallucinations
This data release contains the queries and raw model outputs we analyze in
Dahl et. al, Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis (2024, forthcoming).
Each line represents a query made to an LLM, its response, and an example of a correct response.
This is the public dataset so it does not contains information about all queries made.
Another file, reserve.csv has queries for about… See the full description on the dataset page: https://huggingface.co/datasets/reglab/legal_hallucinations.jspace-hallucination-campaign
jspace hallucination campaign traces (Gemma-4-12B)
24,540 graded Gemma-4-12B responses with internal workspace features, from a
pre-registered cross-domain hallucination detection campaign. Companion to the
jspace repo and the
jspace-lenses model repo
(fitted Jacobian lenses plus the frozen error classifiers trained on this data).
Each row is one prompt, Gemma's greedy answer, an error label, and 31
deployable features read from the model's residual stream during generation
via… See the full description on the dataset page: https://huggingface.co/datasets/solarkyle/jspace-hallucination-campaign.zip-training-hallucination-data-qwen06b-thinking-train-with-valueslegal_hallucinations_paper_datatoolace-unified-hallucinations
ToolACE Unified Hallucination Dataset
This repository contains a unified ToolACE-derived dataset for tool-calling hallucination research.
Files
data/train-00000-of-00001.parquet: leakage-safe grouped training split;
data/test-00000-of-00001.parquet: leakage-safe grouped test split.
The split was rebuilt at the normalized dialogue_id level,the same ToolACE dialogue can't appear in different splits.
Schema
column
description
system
system… See the full description on the dataset page: https://huggingface.co/datasets/VirVen/toolace-unified-hallucinations.toolace-tool-calling-hallucination-ragtruth
ToolACE-derived Tool-Calling Hallucination Dataset
This dataset was created for the course assignment Hallucination Detection in Tool Calling.
It is synthetic by design: starting from ToolACE-style tool-calling dialogues, we automatically inject three required hallucination types:
tool_contradiction
overgeneration
missing_tool
Each example follows a RAGTruth-like format:
query: user query
context: tool output
output: final model answer
hallucination_labels: span-level… See the full description on the dataset page: https://huggingface.co/datasets/marrita/toolace-tool-calling-hallucination-ragtruth.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.Benchmark_Hallucinations_Data
