datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-legal-statutory-hallucination-benchmark
Citation
If you use this dataset, please cite the accompanying paper:
@inproceedings{erdoganyilmaz2026statutoryhallucinations,
title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework},
author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner},
booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)},
year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.brand-hallucination-and-ai-citation-benchmark
🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset
Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3).
📊 Dataset Summary
Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.ai-hallucination-trials
AI Hallucination Trials
4,240 hand-coded trials from a research project on why large language models fabricate confident, false information and what reduces it. In each trial a researcher sent one prompt, usually an invented or obscure acronym or a non-existent institution, to one model under one prompting condition. The card's data holds the prompt, the model's full response, and hand-assigned codes for hallucination and hedging.
Models tested: Gemini, ChatGPT, and Claude.… See the full description on the dataset page: https://huggingface.co/datasets/mahashu/ai-hallucination-trials.hallucination-bert-spans
Hallucination BERT Span Dataset
Flat, one-row-per-span dataset intended for span/token-classification
(BIO-tagging style) hallucination detection over agent tool-calling traces,
derived from the same judging pipeline as the reasoning-distillation set in
this collection.
File
ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join
needed. Each row is one hallucinated span: span (verbatim text), type
(taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.hallucinationThis is a vendored reupload of the Benchmarking Unfaithful Minimal Pairs (BUMP) Dataset available at https://github.com/dataminr-ai/BUMP
The BUMP (Benchmark of Unfaithful Minimal Pairs) dataset stands out as a superior choice for evaluating hallucination detection systems due to its quality and realism. Unlike synthetic datasets such as TruthfulQA, HalluBench, or FaithDial that rely on LLMs to generate hallucinations, BUMP employs human annotators to manually introduce errors into summaries… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/hallucination.toolace-ragtruth-style-hallucinations
ToolACE RAGTruth-style Tool Hallucination Dataset
This dataset was built from ToolACE tool-use dialogues and converted into a RAGTruth-style format for hallucination detection in tool calling.
Task
Given:
query: user query
context: tool response
output: final assistant answer
the goal is to classify whether the answer is grounded in the tool output or belongs to one of three hallucination types.
Labels
clean
tool_output_conflict
overgeneration… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace-ragtruth-style-hallucinations.hallucination_detection_transformers
ToolACE Hallucination Dataset
This dataset was generated for the assignment Hallucination Detection in Tool Calling.
It is based on ToolACE tool-calling dialogues and uses a RAGTruth-style schema:
query: user question
context: tool output / grounding evidence
output: assistant final answer
hallucination_labels: character-level hallucination spans
Files:
File
Rows
Description
toolace_clean_ragtruth.jsonl
1347
clean ToolACE tool-use answers… See the full description on the dataset page: https://huggingface.co/datasets/HASSANI8046/hallucination_detection_transformers.Hallucination_Dataset
