datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rag_hallucinationsProvides examples of hallucinated responses for RAG applications.
audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.HALT_Benchmark_0.1_v1
HALT Benchmark Dataset v1.0
HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse
Overview
HALT is a benchmark for evaluating bounded agentic decision-making under partial observability,
constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity
workflows, where acting too early, failing to escalate, or over-escalating can all be costly.
The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.HalluTruthQA-4K
HalluTruthQA-4K
HalluTruthQA-4K is the official data release for Subtask 2.2 ("Hallucination Detection and Find the Truth") of the HalluScoring 2026 shared task, hosted at ArabicNLP 2026. It extends the HalluTruthQA benchmark from 2,400 to 4,000 expert-annotated Arabic question-answering instances across four knowledge-intensive domains.
Dataset Summary
The full corpus is 4,000 Arabic question-answering instances, exactly balanced across four domains (1,000… See the full description on the dataset page: https://huggingface.co/datasets/Bekhouche/HalluTruthQA-4K.elv-halluc-videos
ELV-Halluc — videos + annotations
A self-contained mirror of the ELV-Halluc benchmark
(CVPR 2026), bundling the raw .mp4 files together with the annotations so the benchmark can be
run without sourcing videos separately.
Paper: arXiv:2508.21496
Original annotations: HLSv/ELV-Halluc (no videos)
Project page: https://elv-halluc.github.io/
This is an unofficial mirror. All credit for the benchmark goes to the original authors; please
cite their paper (below) rather than this… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/elv-halluc-videos.lima-tr
LIMA-TR
This project is dedicated to translating the LIMA (Less Is More for Alignment) dataset from English to Turkish using OpenAI's API (gpt-3.5-turbo)
Source Dataset
LIMA (Less Is More for Alignment) can be found paper-link dataset-link
Korean-Hallucination-Bench
Korean Hallucination Benchmark (한국어 환각 진단 벤치마크)
한국어 LLM의 환각(hallucination) 저항성을 평가하는 4지선다 벤치마크입니다. 법령·특허·행정·의료·금융 5개 전문 도메인에서, 한국어 특화 환각 유형 10종을 다룹니다. 각 문항은 사실 정답 1개와 전문가도 속을 만큼 그럴듯한 환각 오답 3개로 구성됩니다.
통계
총 10,167문항 (4지선다)
도메인(5): 법령 1,901 · 특허 2,190 · 행정 1,864 · 의료 2,113 · 금융 2,099
환각 유형(10): 수치·날짜 오류, 조항 왜곡, 근거 없는 추론, 한자어·신조어 혼재, 과잉 일반화, 멀티턴 맥락 붕괴, 존댓말·반말 역전, 출처 날조, 용어 왜곡, 사실 날조
구축 방법
문항 생성: Darwin-398B-JGOS (문제·선택지)
정답 검수·교정: Claude (Anthropic)… See the full description on the dataset page: https://huggingface.co/datasets/ginigen/Korean-Hallucination-Bench.halluciguard-benchmark
HalluciGuard Benchmark
A 500-sample benchmark for evaluating hallucination detection in Retrieval-Augmented
Generation (RAG) systems, built for the paper HalluciGuard: A Label-Free
Confidence-Aware Hallucination Detection Framework for Retrieval-Augmented Generation
Systems (Anamitra Sarkar).
Construction
Each sample is built from a hand-verified atomic fact rather than downloaded from an
existing QA corpus. Samples are constructed to mirror the query phrasing… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/halluciguard-benchmark.halluciguard-benchmark
HalluciGuard Benchmark
A 500-sample benchmark for evaluating hallucination detection in Retrieval-Augmented
Generation (RAG) systems, built for the paper HalluciGuard: A Label-Free
Confidence-Aware Hallucination Detection Framework for Retrieval-Augmented Generation
Systems.
Construction
Each sample is built from a hand-verified atomic fact rather than downloaded from an
existing QA corpus. Samples are constructed to mirror the query phrasing, difficulty,
and… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/halluciguard-benchmark.
