datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PRE-HAL
PRE-HAL: Multimodal Hallucination Evaluation Benchmark
Dataset Summary
PRE-HAL is a visual question answering (VQA) dataset designed to evaluate and mitigate hallucination in Multimodal Large Language Models (MLLMs). It focuses on testing the model's ability to distinguish between visual perception and parametric knowledge, specifically targeting various hallucination types.
Data Instances
Each instance represents a multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/TerryHWong/PRE-HAL.Med-HALT
Med-HALT: Medical Domain Hallucination Test for Large Language Models
This is a dataset used in the Med-HALT research paper. This research paper focuses on the challenges posed by hallucinations in large language models (LLMs), particularly in the context of the medical domain. We propose a new benchmark and dataset, Med-HALT (Medical Domain Hallucination Test), designed specifically to evaluate hallucinations.
Med-HALT provides a diverse multinational dataset derived from medical… See the full description on the dataset page: https://huggingface.co/datasets/openlifescienceai/Med-HALT.FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.HALP-Bench
HALP-Bench
HALP-Bench is the evaluation benchmark released with the paper
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
(EACL 2026). It aggregates 10,000 image-question pairs drawn from six public sources into a
single, uniformly-formatted set with stable image IDs.
📄 Paper (ACL Anthology): https://aclanthology.org/2026.eacl-long.287/
📄 Preprint (arXiv): https://arxiv.org/abs/2603.05465
💻 Code: https://github.com/Zesearch/HALP… See the full description on the dataset page: https://huggingface.co/datasets/Zesearch/HALP-Bench.Phantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/seyled/Phantom_Hallucination_Detection.whisper-hallucinations
Whisper Hallucinations on Noise
Dataset Summary
This dataset lists common hallucinations from OpenAI Whisper when the input has no speech.
We build it from a noise-only corpus.
We run Whisper on noise clips.
We collect any non-empty text that Whisper outputs.
We deduplicate phrases and count how often they occur.
Use it to test, detect, and reduce non-speech hallucinations.
Motivation
ASR models often output text on silence or noise.
These false hits harm UX… See the full description on the dataset page: https://huggingface.co/datasets/sachaarbonel/whisper-hallucinations.HALAS
Dataset Card for HALAS
Dataset Summary
HALAS (Hallucination Annotations for Large-scale ASR Systems) is a human-annotated dataset of hallucinations produced by modern automatic speech recognition (ASR) systems on real-world speech recordings. The dataset contains span-level hallucination annotations for ASR outputs generated from recordings in the Earnings22 corpus.
HALAS was introduced to address a key limitation in prior hallucination research: most existing… See the full description on the dataset page: https://huggingface.co/datasets/MatBar99/HALAS.narit-ghosts-halo-catalogs
NARIT GHOSTS Halo Catalogs
Dataset Description
This dataset contains reduced stellar catalogs, combined FITS images, and candidate substructure catalogs from the GHOSTS (Galaxy Halos, Outer disks, Substructure, Thick disks, and Star clusters) Survey observed by the Hubble Space Telescope (HST).
It serves as the primary data lake for the automated astronomical pipeline designed to detect faint stellar substructures (like Ultra-Faint Dwarfs and stellar streams) in… See the full description on the dataset page: https://huggingface.co/datasets/appleboiy/narit-ghosts-halo-catalogs.arched-halls-decision-matrix
Arched Halls Decision Matrix / Macierz decyzyjna hal łukowych
Dataset summary
This Polish-language dataset describes 30 practical application scenarios for an arched hall (hala łukowa) in agriculture, storage, logistics, transport, industry, waste management, infrastructure, sports, public facilities, seasonal buildings, construction and energy.
Each record connects the intended use of an arched hall with qualitative decision factors such as indoor climate… See the full description on the dataset page: https://huggingface.co/datasets/halalukowa24/arched-halls-decision-matrix.google-ads-transparencyWebBench
Web Bench: A real-world benchmark for Browser Agents
WebBench is an open, task-oriented benchmark that measures how well browser agents handle realistic web workflows.
It contains 2 ,454 tasks spread across 452 live websites selected from the global top-1000 by traffic.
Last updated: May 28, 2025
Dataset Composition
Category
Description
Example
Count (% of dataset)
READ
Tasks that require searching and extracting information
“Navigate to the news section and… See the full description on the dataset page: https://huggingface.co/datasets/Halluminate/WebBench.BrowserBenchSee: https://github.com/Halluminate/browserbench
song-lyricsHalluciGen
Dataset Card for HalluciGen-Detection
Dataset Summary
This is a dataset for hallucination detection in the paraphrase generation and machine translation scenario. Each example in the dataset consists of a source sentence, a correct hypothesis, and an incorrect hypothesis containing an intrinsic hallucination. A hypothesis is considered to be a hallucination if it is not entailed by the "source" either by containing additional or contradictory information with respect to… See the full description on the dataset page: https://huggingface.co/datasets/NLP-RISE/HalluciGen.legal_rag_hallucinations
Dataset Card for Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools
This data release contains the queries and raw model outputs we analyze in Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies (2024, forthcoming).
Consistent with emerging understanding of AI
benchmarking and leaderboards, we reserve a random sample of 50% of the dataset to… See the full description on the dataset page: https://huggingface.co/datasets/reglab/legal_rag_hallucinations.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/kankshith123/FinQA-hallucination-detection.rag-hallucination-benchmark
RAG Hallucination Benchmark
Context
Retrieval-Augmented Generation (RAG) is the industry standard for reducing LLM hallucinations, but detecting when a RAG system fails is a massive challenge. Most existing benchmarks focus only on massive Deep Learning models and lack tabular features.
This dataset provides a clean, engineered setup to train models (from XGBoost to RoBERTa) to detect hallucinations, predict context faithfulness, and measure answer relevance.… See the full description on the dataset page: https://huggingface.co/datasets/vkshdev/rag-hallucination-benchmark.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
Taming-Hallucinations
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
CVPR 2026 Findings
Project Page | Paper | Code
Dataset Summary
This repository hosts DualityVidQA, the large-scale paired video–QA dataset introduced in
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation.
Taming Hallucinations introduces DualityForge, a controllable diffusion-based framework that turns
real videos into… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/Taming-Hallucinations.Phantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/Anindita1979/Phantom_Hallucination_Detection.HaloQuestThis Dataset introduces challenging pictures for assessing MLLM hallucinations. It includes questions and ground-truth answers.
Original Paper: (https://arxiv.org/abs/2407.15680)[https://arxiv.org/abs/2407.15680]
Code Repo: (https://github.com/google/haloquest/)[https://github.com/google/haloquest/]
halegannada-hosakannada
Halegannada-Hosakannada
I am a Kannada speaker, and I have always wanted to contribute something useful back to the Kannada community as I grow as a researcher. This dataset is just a beginning: a practical, source-backed silver corpus for Halegannada / Old Kannada lexical modernization, source-line glossing, SFT, and preference training.
This became personally meaningful too. My father, Badrinath Gopal, completed the 500-row human source-quality review used in this release. He and… See the full description on the dataset page: https://huggingface.co/datasets/kishanpb/halegannada-hosakannada.userscooking-master-boy-subtitle
Cooking Master Boy Chat Records
Chinese (trditional) subtitle of anime "Cooking Master Boy" (中華一番).
Introduction
This is a collection of subtitles from anime "Cooking Master Boy" (中華一番).
Dataset Description
The dataset is in CSV format, with the following columns:
episode: The episode index of subtitle belogs to.
caption_index: The autoincrement ID of subtitles.
time_start: The starting timecode, which subtitle supposed to appear.
time_end: The ending… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/cooking-master-boy-subtitle.synthetic-imdb-movie-reviews-parallelHalluCounterEval
HalluCounterEval
HalluCounterEval is a large-scale, multi-domain benchmark dataset designed for Reference-Free Hallucination Detection (RFHD) in large language models (LLMs). It supports the evaluation and training of models that detect hallucinated outputs without relying on ground truth answers.
This dataset includes:
Synthetic responses generated by prompting multiple LLMs.
Human-annotated labels for hallucination detection.
Diverse domains: general knowledge, mathematics… See the full description on the dataset page: https://huggingface.co/datasets/ashokurlana/HalluCounterEval.HALT-PROP
About Dataset
The HALT-PROP dataset – the first human-annotated Lithuanian textual propaganda corpus, a pioneering resource not only for Lithuania but also for the broader Baltic region
and other countries neighbouring Russia. The corpus comprises two complementary datasets: (1) 2870 news articles manually labeled by five experts at the article level to identify
the presence of propaganda, and (2) a subset of 1000 articles annotated for specific propaganda techniques and… See the full description on the dataset page: https://huggingface.co/datasets/VilniausUniversitetas/HALT-PROP.HalluciGen-PG
Task 2: HalluciGen - Paraphrase Generation
This dataset contains the trial and test splits per language for the Paraphrase Generation (PG) scenario of the HalluciGen task, which is part of the 2024 ELOQUENT lab.
NOTE: A gold-labeled version of the dataset will be released in a new repository.
Dataset schema
id: unique identifier of the example
source: original model input for paraphrase generation
hyp1: first alternative paraphrase of the source
hyp2: second… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/HalluciGen-PG.
