datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
resultsrequestshallucinations-dporag_hallucinationsProvides examples of hallucinated responses for RAG applications.
ragtruth-translated-hallucinations
RAGTruth Translated Hallucinations
Multilingual machine translation of
RAGTruth into 31 European languages,
preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of
LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the
exact spans that are hallucinated (unsupported by, or contradicting, the provided
context). Here both the RAG prompt and the response are translated into each target
language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.legal_hallucinations_subset
Legal Hallucinations Subset
Dataset Description
This is a curated subset of the reglab/legal_hallucinations dataset, containing up to 1000 randomly sampled rows for each of 6 specific legal reasoning tasks (5444 rows total).
The original dataset was created for the paper: Dahl et al., "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models," Journal of Legal Analysis (2024, forthcoming). Preprint: arxiv:2401.01301
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legal_hallucinations_subset.multi-wiki-qa-synthetic-hallucinationswhisper-hallucinations
Whisper Hallucinations on Noise
Dataset Summary
This dataset lists common hallucinations from OpenAI Whisper when the input has no speech.
We build it from a noise-only corpus.
We run Whisper on noise clips.
We collect any non-empty text that Whisper outputs.
We deduplicate phrases and count how often they occur.
Use it to test, detect, and reduce non-speech hallucinations.
Motivation
ASR models often output text on silence or noise.
These false hits harm UX… See the full description on the dataset page: https://huggingface.co/datasets/sachaarbonel/whisper-hallucinations.legal_hallucinations
Dataset Card for Legal Hallucinations
This data release contains the queries and raw model outputs we analyze in
Dahl et. al, Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis (2024, forthcoming).
Each line represents a query made to an LLM, its response, and an example of a correct response.
This is the public dataset so it does not contains information about all queries made.
Another file, reserve.csv has queries for about… See the full description on the dataset page: https://huggingface.co/datasets/reglab/legal_hallucinations.legal_hallucinations_paper_datatoolace-unified-hallucinations
ToolACE Unified Hallucination Dataset
This repository contains a unified ToolACE-derived dataset for tool-calling hallucination research.
Files
data/train-00000-of-00001.parquet: leakage-safe grouped training split;
data/test-00000-of-00001.parquet: leakage-safe grouped test split.
The split was rebuilt at the normalized dialogue_id level,the same ToolACE dialogue can't appear in different splits.
Schema
column
description
system
system… See the full description on the dataset page: https://huggingface.co/datasets/VirVen/toolace-unified-hallucinations.Benchmark_Hallucinations_DataHallucinations_Benchmark-2legal_rag_hallucinations
Dataset Card for Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools
This data release contains the queries and raw model outputs we analyze in Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies (2024, forthcoming).
Consistent with emerging understanding of AI
benchmarking and leaderboards, we reserve a random sample of 50% of the dataset to… See the full description on the dataset page: https://huggingface.co/datasets/reglab/legal_rag_hallucinations.Taming-Hallucinations
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
CVPR 2026 Findings
Project Page | Paper | Code
Dataset Summary
This repository hosts DualityVidQA, the large-scale paired video–QA dataset introduced in
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation.
Taming Hallucinations introduces DualityForge, a controllable diffusion-based framework that turns
real videos into… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/Taming-Hallucinations.HallucinationsHallucinations_Benchmark-4Hallucinations_Data_Benchmark-3basta_allucinazioni_ai_prompt-enough_hallucinations_ai_prompt
Basta allucinazioni / Enough hallucinations
A small multilingual prompt dataset for AI custom instructions. The goal is to encourage an AI system to admit uncertainty instead of inventing an answer.
Main prompt
Italian original
Se non conosci una risposta ad una domanda che ti è stata fatta non inventare, rispondi semplicemente: non lo so.
English reference
If you do not know the answer to a question you have been asked, do not make… See the full description on the dataset page: https://huggingface.co/datasets/pietrorisipr-2025/basta_allucinazioni_ai_prompt-enough_hallucinations_ai_prompt.Benchmark_Hallucinations_Data__run_Baseer__no_preprocessHallucinations_Data_Benchmark-2wild_hallucinations_with_responsesgrounded-vs-fabricated-hallucinations
Grounded vs. Fabricated Hallucinations
This dataset consists of hallucinated and grounded answers to the first 3000 rows of TriviaQA rc.nocontext validation split.
Methodology
The dataset consists of a training, evaluation, and test split. Truthful and hallucinated answers overlap in the same window, so for every truthful answer there is at least
one corresponding hallucinated answer. Hallucinated answers are not organic but rather directly prompted for via gaslighting in… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/grounded-vs-fabricated-hallucinations.Hallucinations_BenchmarkHallucinations_Benchmark-3Hallucinations_Data_BenchmarkBenchmark_Hallucinations_Data__run_Baseer__grayscaleBenchmark_Hallucinations_Data__run_Baseer__grayscale__remove_borderHallucinations_Benchmark-4__run_Baseer__webpBenchmark_Hallucinations_Data__run_Baseer__denoise_nlm
