CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01satpalsr /hallucinations-dpotext1K<n<10K5 likes905 downloads3y agoHugging Face02aporia-ai /rag_hallucinationsProvides examples of hallucinated responses for RAG applications. textquestion-answering1K<n<10K9 likes453 downloads2y agoHugging Face03Mozilla /link_tab_hallucination_eval link_tab_hallucination_eval Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns (false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts. Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result (a frozen page snapshot) are baked into the message thread so predictions are reproducible (no live fetch), while the final scorable user turn still shows the real tab URL. 139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.textn<1K0 likes288 downloads2mo agoHugging Face04aseth125 /audio-hallucination-attack Audio Hallucination Attacks (AHA) Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models" It contains two subsets: AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training Audio Files The audio files are provided as compressed archives in this repository: File Contents Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.audioaudio-classification100K<n<1M2 likes191 downloads6mo agoHugging Face05Groundtruth-Data /groundtruth-hallucination-bench-sample Groundtruth Data Hallucination Benchmark Sample This public teaser contains 180 representative, source-backed examples from Groundtruth Data products. Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is: Find where a model fails. Prove the failure with a larger verified evaluation. Provide targeted remediation/training data. Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.textquestion-answeringn<1K1 likes82 downloads23h agoHugging Face06LawChatAI /turkish-legal-statutory-hallucination-benchmark Citation If you use this dataset, please cite the accompanying paper: @inproceedings{erdoganyilmaz2026statutoryhallucinations, title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework}, author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner}, booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)}, year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.tabularn<1K2 likes74 downloads4mo agoHugging Face07stindardlogic /hallucination-reduction-dpo-100k Hallucination Reduction DPO (100K) 100,000 DPO preference pairs training LLMs to stay within knowledge bounds. The chosen response is accurate and appropriately uncertain; the rejected response is confident but wrong — fabricated statistics, fake citations, wrong facts, overclaimed certainty. Motivation Hallucination is the #1 reliability concern blocking enterprise LLM adoption. Models fail in predictable patterns: Inventing specific statistics with false… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-reduction-dpo-100k.texttext-generation100K<n<1M0 likes55 downloads2mo agoHugging Face08ssurface /hallucination-bert-spans Hallucination BERT Span Dataset Flat, one-row-per-span dataset intended for span/token-classification (BIO-tagging style) hallucination detection over agent tool-calling traces, derived from the same judging pipeline as the reasoning-distillation set in this collection. File ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join needed. Each row is one hallucinated span: span (verbatim text), type (taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.tabular10K<n<100K0 likes47 downloads2mo agoHugging Face09pixeloffice /brand-hallucination-and-ai-citation-benchmark 🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3). 📊 Dataset Summary Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.tabulartext-classificationn<1K0 likes44 downloads1mo agoHugging Face10mahashu /ai-hallucination-trials AI Hallucination Trials 4,240 hand-coded trials from a research project on why large language models fabricate confident, false information and what reduces it. In each trial a researcher sent one prompt, usually an invented or obscure acronym or a non-existent institution, to one model under one prompting condition. The card's data holds the prompt, the model's full response, and hand-assigned codes for hallucination and hedging. Models tested: Gemini, ChatGPT, and Claude.… See the full description on the dataset page: https://huggingface.co/datasets/mahashu/ai-hallucination-trials.tabulartext-classification1K<n<10K0 likes44 downloads5d agoHugging Face11HassanB4 /sawb-arabic-hallucination-dataset Sawb Glossary Hallucination Examples Overview This dataset contains 158 synthesized examples of cultural hallucination in Arabic AI terminology, built as part of Sawb, a detect-then-explain system for cultural hallucination detection in Arabic LLM outputs, developed for the ICAIRE 2026 Hackathon Track 3 (First Place, Cultural Hallucination Tools). Each record captures a case where the DeepSeek API was asked to define an AI/ML term in Arabic without any grounding… See the full description on the dataset page: https://huggingface.co/datasets/HassanB4/sawb-arabic-hallucination-dataset.texttext-classificationn<1K0 likes40 downloads26d agoHugging Face12GuardrailsAI /hallucinationThis is a vendored reupload of the Benchmarking Unfaithful Minimal Pairs (BUMP) Dataset available at https://github.com/dataminr-ai/BUMP The BUMP (Benchmark of Unfaithful Minimal Pairs) dataset stands out as a superior choice for evaluating hallucination detection systems due to its quality and realism. Unlike synthetic datasets such as TruthfulQA, HalluBench, or FaithDial that rely on LLMs to generate hallucinations, BUMP employs human annotators to manually introduce errors into summaries… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/hallucination.tabulartext-classificationn<1K2 likes37 downloads2y agoHugging Face13ginigen /Korean-Hallucination-Bench Korean Hallucination Benchmark (한국어 환각 진단 벤치마크) 한국어 LLM의 환각(hallucination) 저항성을 평가하는 4지선다 벤치마크입니다. 법령·특허·행정·의료·금융 5개 전문 도메인에서, 한국어 특화 환각 유형 10종을 다룹니다. 각 문항은 사실 정답 1개와 전문가도 속을 만큼 그럴듯한 환각 오답 3개로 구성됩니다. 통계 총 10,167문항 (4지선다) 도메인(5): 법령 1,901 · 특허 2,190 · 행정 1,864 · 의료 2,113 · 금융 2,099 환각 유형(10): 수치·날짜 오류, 조항 왜곡, 근거 없는 추론, 한자어·신조어 혼재, 과잉 일반화, 멀티턴 맥락 붕괴, 존댓말·반말 역전, 출처 날조, 용어 왜곡, 사실 날조 구축 방법 문항 생성: Darwin-398B-JGOS (문제·선택지) 정답 검수·교정: Claude (Anthropic)… See the full description on the dataset page: https://huggingface.co/datasets/ginigen/Korean-Hallucination-Bench.textquestion-answering10K<n<100K3 likes35 downloads3mo agoHugging Face14rohit901 /nlp_proj_llm_hallucinationtextn<1K0 likes34 downloads3y agoHugging Face15dzur658 /grounded-vs-fabricated-hallucinations Grounded vs. Fabricated Hallucinations This dataset consists of hallucinated and grounded answers to the first 3000 rows of TriviaQA rc.nocontext validation split. Methodology The dataset consists of a training, evaluation, and test split. Truthful and hallucinated answers overlap in the same window, so for every truthful answer there is at least one corresponding hallucinated answer. Hallucinated answers are not organic but rather directly prompted for via gaslighting in… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/grounded-vs-fabricated-hallucinations.text1K<n<10K0 likes34 downloads6mo agoHugging Face16stindardlogic /hallucination-grounding-dpo-4k Hallucination Grounding DPO Pairs (4K) DPO preference pairs targeting the full spectrum of factuality failures — from hallucination to over-hedging. Motivation Existing refusal/safety datasets focus on what not to say. This dataset targets the orthogonal challenge: when to say "I don't know" vs. when to answer confidently. Models that over-refuse waste user trust; models that hallucinate destroy it. Dataset Description 4,000 preference pairs across… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-grounding-dpo-4k.texttext-generation1K<n<10K0 likes33 downloads2mo agoHugging Face17Debarun12 /cybersec-hallucination-guard-dataset dataset_card_content = """--- license: apache-2.0 task_categories: - question-answering - text-classification tags: - cybersecurity - hallucination-detection - rag - groundedness - synthetic language: - en size_categories: - 1K<n<10K Cybersecurity Hallucination Detection Dataset This dataset was built to train Cybersec Hallucination Guard, a LoRA-tuned model that detects whether a retrieved context contains enough information to answer a given… See the full description on the dataset page: https://huggingface.co/datasets/Debarun12/cybersec-hallucination-guard-dataset.text1K<n<10K1 likes32 downloads2mo agoHugging Face18Ali-Bhai /toolace-ragtruth-style-hallucinations ToolACE RAGTruth-style Tool Hallucination Dataset This dataset was built from ToolACE tool-use dialogues and converted into a RAGTruth-style format for hallucination detection in tool calling. Task Given: query: user query context: tool response output: final assistant answer the goal is to classify whether the answer is grounded in the tool output or belongs to one of three hallucination types. Labels clean tool_output_conflict overgeneration… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace-ragtruth-style-hallucinations.tabulartext-classification1K<n<10K1 likes27 downloads4mo agoHugging Face19SamarJaffri /hallucination-groundedness-blindspottextn<1K0 likes14 downloads18h agoHugging Face20abramov-de /toolace-hallucination-spans ToolACE Hallucination Spans (Assignment 3) This dataset is constructed from Team-ACE/ToolACE by injecting three hallucination types with span-level labels: contradiction overgeneration missing_tool Schema (RAGTruth-compatible): query, context, output, hallucination_labels. texttext-classificationn<1K0 likes12 downloads4mo agoHugging Face21ssurface /hallucination-distill-reasoninggated Hallucination Reasoning Distillation Set Chain-of-thought reasoning + consensus-voted hallucination spans, distilled from gpt-oss-120b acting as a judge over real agent tool-calling traces sampled from Agent-Ark/Toucan-1.5M. Used to LoRA-SFT a small Qwen model (see qwen_distill_sft/) to reproduce the teacher's reasoning + span output on unseen traces. Files ds_thinking_traces_with_messages.jsonl — use this one. 2,360 traces, one row per trace, self-contained: the… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-distill-reasoning.text1K<n<10K0 likes11 downloads2mo agoHugging Face22mukunda1729 /hallucination-risk-cases hallucination-risk-cases 20 hand-labeled (prompt → response → ground-truth) tuples covering common LLM hallucination failure modes. Each case is rated for hallucination risk so you can evaluate whether your detector / scorer / judge correctly distinguishes the safe responses from the fabricated ones. Categories Category Count What it tests factual 4 Straightforward verifiable facts fabricated-citation 1 Invented academic citations fabricated-api 1 Invented… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/hallucination-risk-cases.textn<1K0 likes10 downloads5mo agoHugging Face23drond0174 /RAGTruth-Hallucinations ToolACE ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. More details… See the full description on the dataset page: https://huggingface.co/datasets/drond0174/RAGTruth-Hallucinations.texttext-generation10K<n<100K0 likes10 downloads4mo agoHugging Face24SwarmandBee /defendable-pain-hallucination-pain-v0.1 Pain Receipt · Hallucination Pain · v0.1 Watchlist "To the shed. Honest about the gaps." — Mr. Defendable This pain mode is NOT yet receipt-anchored in the Defendable v0.1 corpus. Rather than fabricate rows, we publish a watchlist — explicit operator notes about where this receipt class will land in v0.2. Part of the SwarmandBee 100-pack. All 100 are free · CC-BY-4.0 · all honest about what they are and aren't. Why a watchlist instead of fabricated rows No proof, no… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-hallucination-pain-v0.1.texttext-classificationn<1K0 likes9 downloads4mo agoHugging Face25linweitao /dllm-hallucination dllm-hallucination Precomputed trajectories and baselines for the NeurIPS 2026 paper Trajectory Unified Entropy (TUE): A Training-Free Hallucination Detector for Masked Diffusion Language Models. Contents Trajectories (trajectories/) Per-sample denoising trajectory uncertainty maps for 3 models × 6 datasets × 2 sampling depths (T=64, T=128). Schema per traj_{model}_{task}_T{T}_L64{_judged}.npz: entropies — float16, shape (N=400, T, L=64); token-distribution… See the full description on the dataset page: https://huggingface.co/datasets/linweitao/dllm-hallucination.textn<1K0 likes8 downloads5mo agoHugging Face26ka3094 /dpo_hallucination_llama3-8bittext1K<n<10K0 likes7 downloads10mo agoHugging Face27ranazayed19 /Arabic_Hallucinationtextn<1K0 likes7 downloads8mo agoHugging Face28scale-safety-research /internet_capability_hallucinationA dataset to evaluate if models are more likely to be mistaken about their capabilities if the task is morally questionable. Each entry has an action that requires using the internet (e.g. "Do a Google search"), a specific prompt that asks the model to do the action for some "nice" purpose, and a prompt that asks the model to do it for a slightly-evil purpose. The hypothesis to test here is that some models are likely to hallucinate their internet capabilities and say "I've searched google and… See the full description on the dataset page: https://huggingface.co/datasets/scale-safety-research/internet_capability_hallucination.textn<1K0 likes6 downloads2y agoHugging Face29Jahanshahi /elm-arabic-hallucination-beacons 🍄 Arabic LLM Hallucination Dataset - Regional Comparisons Team Beacons (المنارات) | ELM NLP Challenge | MenaML Winter School 2026 📋 Dataset Description This dataset contains 10,000 Arabic prompts designed to trigger hallucinations in Large Language Models by exploiting their tendency to fabricate non-existent regional differences in Saudi Arabian culture. Why This Works When asked "What's the difference between X in Region1 vs Region2?", LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Jahanshahi/elm-arabic-hallucination-beacons.text10K<n<100K0 likes6 downloads8mo agoHugging Face30Gretfel /hallucination-prompts-in-Arabic-LLMstext10K<n<100K0 likes6 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.