datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrowseComp-V3
BrowseComp-V3: A Benchmark Dataset for Multimodal Browsing Agents
A dataset containing 300 samples with encrypted question-answer pairs, images, search trajectories, and sub-goals.
Contents
├── data/
│ ├── train.jsonl # Main dataset (1.44 MB, 300 samples)
│ └── images/ # Referenced images
├── scripts/
│ ├── decryption_script.py # Decrypt entire dataset
│ ├── decrypt_batch.py # Batch decrypt to files
│ ├── encryption_utils.py… See the full description on the dataset page: https://huggingface.co/datasets/Halcyon-Zhang/BrowseComp-V3.hallucinations-dporag_hallucinationsProvides examples of hallucinated responses for RAG applications.
link_tab_hallucination_eval
link_tab_hallucination_eval
Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns
(false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts.
Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result
(a frozen page snapshot) are baked into the message thread so predictions are reproducible
(no live fetch), while the final scorable user turn still shows the real tab URL.
139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.Vript-HAL
🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo]
Vript-HAL
Vript-HAL is the first benchmark evaluating action and object hallucinations in video LLMs
Getting Started
By downloading these datasets, you agree to the terms of the License.
Vript-HAL/
|
├── HAL_scenes/
│ ├── -_MRAAhEKio-Scene-010.mp4
│ └── ...
│
└── HAL_annotations.jsonl
HAL_scenes: The trimmed video clips in the Vript-HAL benchmark.
HAL_annotations.jsonl: The… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript-HAL.HalluToolACEaudio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.Adala_Fonction_Publique_Maroc_Arabic_json_dataset
🏛️ Adala Fonction Publique Maroc Arabic Dataset
This dataset contains structured legal data extracted from Moroccan public service law texts, sourced from adala.justice.gov.ma. The content is in Arabic and is designed to support NLP and AI applications in legal tech, especially for Moroccan administrative and public law.
📂 Dataset Structure
Format: JSON
Language: Arabic (Standard & Legal dialect)
Content: Articles, chapters, titles from Moroccan public law texts… See the full description on the dataset page: https://huggingface.co/datasets/halimbahae/Adala_Fonction_Publique_Maroc_Arabic_json_dataset.HalluTruthQA-4K
HalluTruthQA-4K
HalluTruthQA-4K is the official data release for Subtask 2.2 ("Hallucination Detection and Find the Truth") of the HalluScoring 2026 shared task, hosted at ArabicNLP 2026. It extends the HalluTruthQA benchmark from 2,400 to 4,000 expert-annotated Arabic question-answering instances across four knowledge-intensive domains.
Dataset Summary
The full corpus is 4,000 Arabic question-answering instances, exactly balanced across four domains (1,000… See the full description on the dataset page: https://huggingface.co/datasets/Bekhouche/HalluTruthQA-4K.TDC_half_life_obachwikiMIA-2024-hard
WikiMIA-2024 Hard Dataset
Dataset Description
WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs.
This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques.
It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.HalluScope-30K
HalluScope-30K
HalluScope-30K is a large-scale dataset for fine-grained hallucination
diagnosis in multimodal large language models (MLLMs). Each sample pairs an
image with a model-generated response in which every hallucinated span is
annotated with one of 12 fine-grained hallucination types.
<Tagged_Text>
The <hallucination type="Color_Attribute">bright red</hallucination>
<hallucination type="Object">sports</hallucination> car is
<hallucination type="Spatial_Attribute">parked… See the full description on the dataset page: https://huggingface.co/datasets/wkinglin/HalluScope-30K.clustering-hal-s2s
Clustering HAL
This dataset was created by scrapping data from the HAL platform.
Over 80,000 articles have been scrapped to keep their id, title and category.
It was originally used for the French version of MTEB, but it can also be used for various clustering or classification tasks, or even evaluate the general knowledge of a model.
⚠️ This dataset contains 2 subsets. IT IS STRONGLY ADVISED TO USE THE CLEANED UP mteb_eval SUBSET:
"raw" subset : contains the data originally… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/clustering-hal-s2s.dllm-effect-parents-w128-tau2-half-v2
Deprecated — do not use
The payload on this repository's main branch was removed on 2026-09-19.
It used the superseded pre-fix counterfactual dependency construction
(manifest.json SHA-256 678fc4069de11c941120e3bfe431857ea15d70b8a02bde34a1e5091ad82f0888) and is not the corrected
d1-marginal graph.
Use the corrected canonical B64 replacement:
zimplex/dllm-effect-parents-llada2-finemath-half-d1marginal-w128-tau2-b64-v3
(manifest… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-effect-parents-w128-tau2-half-v2.HALT_Benchmark_0.1_v1
HALT Benchmark Dataset v1.0
HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse
Overview
HALT is a benchmark for evaluating bounded agentic decision-making under partial observability,
constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity
workflows, where acting too early, failing to escalate, or over-escalating can all be costly.
The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.Hala-4.6M-SFT
Hala: Arabic-Centric Instruction & Translation Dataset
Paper: Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
Authors: Hasan Abed Al Kader Hammoud*, Mohammad Zbeeb*, Bernard Ghanem
Affiliation: King Abdullah University of Science and Technology (KAUST)
*Equal contribution
In Arabic, حلا (Hala) conveys sweetness and beauty—qualities long associated with the language itself. In this spirit, we extend Hala to datasets that aim to enrich… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/Hala-4.6M-SFT.HALO-Gemini-3-Flash-AppWorld
Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal)
Dataset Overview
This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem.
Field
Value
Model
Gemini 3 Flash
Benchmark
AppWorld
Split
test-normal
Total Traces
168
Total Spans
3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.AuthorMix
[StyleRemix] AuthorMix Dataset
Dataset Description
This contains the AuthorMix dataset, which is created for authorship obfuscation. It includes data from four distinct domains: presidential speeches, early-1900s fiction novels, scholarly articles, and diary-style blogs. Altogether, AuthorMix contains over 30k high-quality paragraphs from 14 authors.
This work was created in the paper: StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/AuthorMix.HalluCompass
HalluCompass
A direction-aware diagnostic benchmark and protocol for vision-language model (VLM) hallucination evaluation
NeurIPS 2026 Datasets & Benchmarks Track
Unified release: 2,200 images across MS-COCO + AMBER + NoCaps + VizWiz, 10,000 queries, two-pass annotation (GPT-4o-mini 1st-pass + 5-human verification, Fleiss' κ = 0.72 inter-annotator agreement on a stratified 250-image validation subset, substantial agreement per Landis & Koch 1977). A balanced POPE-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymous80934/HalluCompass.HalluciGen-Translation
Task 2: HalluciGen - Tranlsation
This dataset contains the trial and test splits per language pair for the Translation scenario of the HalluciGen task, which is part of the 2024 ELOQUENT lab.
NOTE: A gold-labeled version of the dataset will be released in a new repository.
Dataset schema
id: unique identifier of the example
langpair: the source and target language pair of the example
source: original model input for translation
hyp1: first alternative translation of the… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/HalluciGen-Translation.HalloMTBench
HalloMTBench: A Benchmark for Translation Hallucination in LLMs
Paper | GitHub
Dataset Summary
HalloMTBench is a new and challenging benchmark designed to evaluate the performance of Large Language Models (LLMs) against translation hallucinations.
The result is a high-quality, expert-verified dataset of 6,908 challenging samples that capture naturally occurring hallucinations, providing a cost-effective and robust tool for evaluating model safety and… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/HalloMTBench.second_half_trainingturkish-legal-statutory-hallucination-benchmark
Citation
If you use this dataset, please cite the accompanying paper:
@inproceedings{erdoganyilmaz2026statutoryhallucinations,
title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework},
author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner},
booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)},
year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.ELV-Halluc-DPO
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
[📖 arXiv Paper] [🤗 Dataset] [🐙 code]
ELV-Halluc is designed for long-video hallucination evaluation, especially enables a systematic investigation of SAH(Semantic Aggregation Hallucinations).
👀 ELV-Halluc Overview
ELV-Halluc contains 4,800 binary QA pairs, which can be grouped into 3,200 adversarial QA pairs.
For each selected video, we construct 24 binary QA pairs by… See the full description on the dataset page: https://huggingface.co/datasets/HLSv/ELV-Halluc-DPO.halo-guard-bench
HALO Guard Bench
A constitutional, multilingual benchmark and training corpus for LLM input safety classification.
Built by Astroware · Released June 2026
Why another safety benchmark?
Every existing public safety benchmark has the same structural flaw: it was designed to measure the wrong thing.
WildGuard, ToxicChat, Aegis, HarmBench, and OpenAI Moderation all share a common architecture — human annotators (or a prompted model) label a stream of observed chat… See the full description on the dataset page: https://huggingface.co/datasets/astroware/halo-guard-bench.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.chore-chart-kit
chore-chart-kit
100 open-source printable chore chart templates, hand-curated, CC-BY-4.0.
This dataset is the canonical machine-readable index of the
chore-chart-kit project — a free,
open-source collection of printable chore-chart designs for parents, teachers, and
makers. Each row describes one template: theme, format, color palette, descriptive
copy, and URLs to the JSON template, the rendered PNG/SVG/PDF samples, and the
matching customizable web editor on… See the full description on the dataset page: https://huggingface.co/datasets/halallens-no/chore-chart-kit.hallucination-reduction-dpo-100k
Hallucination Reduction DPO (100K)
100,000 DPO preference pairs training LLMs to stay within knowledge bounds. The chosen response is accurate and appropriately uncertain; the rejected response is confident but wrong — fabricated statistics, fake citations, wrong facts, overclaimed certainty.
Motivation
Hallucination is the #1 reliability concern blocking enterprise LLM adoption. Models fail in predictable patterns:
Inventing specific statistics with false… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-reduction-dpo-100k.hallucination-bert-spans
Hallucination BERT Span Dataset
Flat, one-row-per-span dataset intended for span/token-classification
(BIO-tagging style) hallucination detection over agent tool-calling traces,
derived from the same judging pipeline as the reasoning-distillation set in
this collection.
File
ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join
needed. Each row is one hallucinated span: span (verbatim text), type
(taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.plm-hallubench
PLM-HalluBench: A Multi-level Benchmark for Evaluating Hallucinations in Protein Language Models
NeurIPS 2026 Evaluations and Datasets Track submission (double-blind).
PLM-HalluBench is a benchmark for evaluating hallucination in protein language models (PLMs) — outputs that look like proteins but violate basic biophysics. It is organised around a three-level taxonomy that separates sequence-, structure-, and function-level failure modes, and it pairs a Factual track (BPHS against a… See the full description on the dataset page: https://huggingface.co/datasets/plm-hallubench/plm-hallubench.
