CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Halcyon-Zhang /BrowseComp-V3 BrowseComp-V3: A Benchmark Dataset for Multimodal Browsing Agents A dataset containing 300 samples with encrypted question-answer pairs, images, search trajectories, and sub-goals. Contents ├── data/ │ ├── train.jsonl # Main dataset (1.44 MB, 300 samples) │ └── images/ # Referenced images ├── scripts/ │ ├── decryption_script.py # Decrypt entire dataset │ ├── decrypt_batch.py # Batch decrypt to files │ ├── encryption_utils.py… See the full description on the dataset page: https://huggingface.co/datasets/Halcyon-Zhang/BrowseComp-V3.imagen<1K6 likes3.9k downloads7mo agoHugging Face02satpalsr /hallucinations-dpotext1K<n<10K5 likes908 downloads3y agoHugging Face03aporia-ai /rag_hallucinationsProvides examples of hallucinated responses for RAG applications. textquestion-answering1K<n<10K9 likes464 downloads2y agoHugging Face04Mozilla /link_tab_hallucination_eval link_tab_hallucination_eval Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns (false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts. Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result (a frozen page snapshot) are baked into the message thread so predictions are reproducible (no live fetch), while the final scorable user turn still shows the real tab URL. 139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.textn<1K0 likes304 downloads2mo agoHugging Face05Mutonix /Vript-HAL 🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] Vript-HAL Vript-HAL is the first benchmark evaluating action and object hallucinations in video LLMs Getting Started By downloading these datasets, you agree to the terms of the License. Vript-HAL/ | ├── HAL_scenes/ │ ├── -_MRAAhEKio-Scene-010.mp4 │ └── ... │ └── HAL_annotations.jsonl HAL_scenes: The trimmed video clips in the Vript-HAL benchmark. HAL_annotations.jsonl: The… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript-HAL.textvideo-classificationn<1K5 likes285 downloads2y agoHugging Face06annnettte /HalluToolACEtext1K<n<10K0 likes235 downloads4mo agoHugging Face07aseth125 /audio-hallucination-attack Audio Hallucination Attacks (AHA) Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models" It contains two subsets: AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training Audio Files The audio files are provided as compressed archives in this repository: File Contents Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.audioaudio-classification100K<n<1M2 likes197 downloads6mo agoHugging Face08halimbahae /Adala_Fonction_Publique_Maroc_Arabic_json_dataset 🏛️ Adala Fonction Publique Maroc Arabic Dataset This dataset contains structured legal data extracted from Moroccan public service law texts, sourced from adala.justice.gov.ma. The content is in Arabic and is designed to support NLP and AI applications in legal tech, especially for Moroccan administrative and public law. 📂 Dataset Structure Format: JSON Language: Arabic (Standard & Legal dialect) Content: Articles, chapters, titles from Moroccan public law texts… See the full description on the dataset page: https://huggingface.co/datasets/halimbahae/Adala_Fonction_Publique_Maroc_Arabic_json_dataset.textn<1K0 likes186 downloads1y agoHugging Face09Bekhouche /HalluTruthQA-4K HalluTruthQA-4K HalluTruthQA-4K is the official data release for Subtask 2.2 ("Hallucination Detection and Find the Truth") of the HalluScoring 2026 shared task, hosted at ArabicNLP 2026. It extends the HalluTruthQA benchmark from 2,400 to 4,000 expert-annotated Arabic question-answering instances across four knowledge-intensive domains. Dataset Summary The full corpus is 4,000 Arabic question-answering instances, exactly balanced across four domains (1,000… See the full description on the dataset page: https://huggingface.co/datasets/Bekhouche/HalluTruthQA-4K.texttext-classification1K<n<10K0 likes152 downloads28d agoHugging Face10scikit-fingerprints /TDC_half_life_obachn<1K0 likes151 downloads2y agoHugging Face11hallisky /wikiMIA-2024-hard WikiMIA-2024 Hard Dataset Dataset Description WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs. This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques. It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.tabulartext-classification1K<n<10K0 likes138 downloads1y agoHugging Face12wkinglin /HalluScope-30K HalluScope-30K HalluScope-30K is a large-scale dataset for fine-grained hallucination diagnosis in multimodal large language models (MLLMs). Each sample pairs an image with a model-generated response in which every hallucinated span is annotated with one of 12 fine-grained hallucination types. <Tagged_Text> The <hallucination type="Color_Attribute">bright red</hallucination> <hallucination type="Object">sports</hallucination> car is <hallucination type="Spatial_Attribute">parked… See the full description on the dataset page: https://huggingface.co/datasets/wkinglin/HalluScope-30K.imagevisual-question-answeringn<1K0 likes111 downloads2mo agoHugging Face13lyon-nlp /clustering-hal-s2s Clustering HAL This dataset was created by scrapping data from the HAL platform. Over 80,000 articles have been scrapped to keep their id, title and category. It was originally used for the French version of MTEB, but it can also be used for various clustering or classification tasks, or even evaluate the general knowledge of a model. ⚠️ This dataset contains 2 subsets. IT IS STRONGLY ADVISED TO USE THE CLEANED UP mteb_eval SUBSET: "raw" subset : contains the data originally… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/clustering-hal-s2s.texttext-classification100K<n<1M1 likes106 downloads2y agoHugging Face14zimplex /dllm-effect-parents-w128-tau2-half-v2 Deprecated — do not use The payload on this repository's main branch was removed on 2026-09-19. It used the superseded pre-fix counterfactual dependency construction (manifest.json SHA-256 678fc4069de11c941120e3bfe431857ea15d70b8a02bde34a1e5091ad82f0888) and is not the corrected d1-marginal graph. Use the corrected canonical B64 replacement: zimplex/dllm-effect-parents-llada2-finemath-half-d1marginal-w128-tau2-b64-v3 (manifest… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-effect-parents-w128-tau2-half-v2.textn<1K0 likes106 downloads3d agoHugging Face15supreme-lab /HALT_Benchmark_0.1_v1 HALT Benchmark Dataset v1.0 HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse Overview HALT is a benchmark for evaluating bounded agentic decision-making under partial observability, constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity workflows, where acting too early, failing to escalate, or over-escalating can all be costly. The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.texttext-classification1K<n<10K1 likes105 downloads5mo agoHugging Face16hammh0a /Hala-4.6M-SFT Hala: Arabic-Centric Instruction & Translation Dataset Paper: Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale Authors: Hasan Abed Al Kader Hammoud*, Mohammad Zbeeb*, Bernard Ghanem Affiliation: King Abdullah University of Science and Technology (KAUST) *Equal contribution In Arabic, حلا (Hala) conveys sweetness and beauty—qualities long associated with the language itself. In this spirit, we extend Hala to datasets that aim to enrich… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/Hala-4.6M-SFT.texttext-generation1M<n<10M5 likes103 downloads1y agoHugging Face17inference-net /HALO-Gemini-3-Flash-AppWorld Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal) Dataset Overview This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem. Field Value Model Gemini 3 Flash Benchmark AppWorld Split test-normal Total Traces 168 Total Spans 3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.texttext-generation1K<n<10K7 likes103 downloads5mo agoHugging Face18hallisky /AuthorMix [StyleRemix] AuthorMix Dataset Dataset Description This contains the AuthorMix dataset, which is created for authorship obfuscation. It includes data from four distinct domains: presidential speeches, early-1900s fiction novels, scholarly articles, and diary-style blogs. Altogether, AuthorMix contains over 30k high-quality paragraphs from 14 authors. This work was created in the paper: StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/AuthorMix.texttext-classification10K<n<100K5 likes97 downloads2y agoHugging Face19anonymous80934 /HalluCompass HalluCompass A direction-aware diagnostic benchmark and protocol for vision-language model (VLM) hallucination evaluation NeurIPS 2026 Datasets & Benchmarks Track Unified release: 2,200 images across MS-COCO + AMBER + NoCaps + VizWiz, 10,000 queries, two-pass annotation (GPT-4o-mini 1st-pass + 5-human verification, Fleiss' κ = 0.72 inter-annotator agreement on a stratified 250-image validation subset, substantial agreement per Landis & Koch 1977). A balanced POPE-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymous80934/HalluCompass.imagevisual-question-answering10K<n<100K0 likes91 downloads5mo agoHugging Face20Eloquent /HalluciGen-Translation Task 2: HalluciGen - Tranlsation This dataset contains the trial and test splits per language pair for the Translation scenario of the HalluciGen task, which is part of the 2024 ELOQUENT lab. NOTE: A gold-labeled version of the dataset will be released in a new repository. Dataset schema id: unique identifier of the example langpair: the source and target language pair of the example source: original model input for translation hyp1: first alternative translation of the… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/HalluciGen-Translation.text1K<n<10K0 likes83 downloads2y agoHugging Face21ATH-MaaS /HalloMTBench HalloMTBench: A Benchmark for Translation Hallucination in LLMs Paper | GitHub Dataset Summary HalloMTBench is a new and challenging benchmark designed to evaluate the performance of Large Language Models (LLMs) against translation hallucinations. The result is a high-quality, expert-verified dataset of 6,908 challenging samples that capture naturally occurring hallucinations, providing a cost-effective and robust tool for evaluating model safety and… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/HalloMTBench.texttranslation1K<n<10K7 likes81 downloads8mo agoHugging Face22timhua /second_half_trainingtext10K<n<100K1 likes79 downloads1y agoHugging Face23LawChatAI /turkish-legal-statutory-hallucination-benchmark Citation If you use this dataset, please cite the accompanying paper: @inproceedings{erdoganyilmaz2026statutoryhallucinations, title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework}, author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner}, booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)}, year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.tabularn<1K2 likes71 downloads4mo agoHugging Face24HLSv /ELV-Halluc-DPO ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding [📖 arXiv Paper] [🤗 Dataset] [🐙 code] ELV-Halluc is designed for long-video hallucination evaluation, especially enables a systematic investigation of SAH(Semantic Aggregation Hallucinations). 👀 ELV-Halluc Overview ELV-Halluc contains 4,800 binary QA pairs, which can be grouped into 3,200 adversarial QA pairs. For each selected video, we construct 24 binary QA pairs by… See the full description on the dataset page: https://huggingface.co/datasets/HLSv/ELV-Halluc-DPO.textvideo-text-to-text1K<n<10K1 likes64 downloads1y agoHugging Face25astroware /halo-guard-bench HALO Guard Bench A constitutional, multilingual benchmark and training corpus for LLM input safety classification. Built by Astroware · Released June 2026 Why another safety benchmark? Every existing public safety benchmark has the same structural flaw: it was designed to measure the wrong thing. WildGuard, ToxicChat, Aegis, HarmBench, and OpenAI Moderation all share a common architecture — human annotators (or a prompted model) label a stream of observed chat… See the full description on the dataset page: https://huggingface.co/datasets/astroware/halo-guard-bench.texttext-classification10K<n<100K0 likes60 downloads3mo agoHugging Face26Groundtruth-Data /groundtruth-hallucination-bench-sample Groundtruth Data Hallucination Benchmark Sample This public teaser contains 180 representative, source-backed examples from Groundtruth Data products. Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is: Find where a model fails. Prove the failure with a larger verified evaluation. Provide targeted remediation/training data. Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.textquestion-answeringn<1K0 likes59 downloads23d agoHugging Face27halallens-no /chore-chart-kit chore-chart-kit 100 open-source printable chore chart templates, hand-curated, CC-BY-4.0. This dataset is the canonical machine-readable index of the chore-chart-kit project — a free, open-source collection of printable chore-chart designs for parents, teachers, and makers. Each row describes one template: theme, format, color palette, descriptive copy, and URLs to the JSON template, the rendered PNG/SVG/PDF samples, and the matching customizable web editor on… See the full description on the dataset page: https://huggingface.co/datasets/halallens-no/chore-chart-kit.imagetabular-classificationn<1K1 likes58 downloads4mo agoHugging Face28stindardlogic /hallucination-reduction-dpo-100k Hallucination Reduction DPO (100K) 100,000 DPO preference pairs training LLMs to stay within knowledge bounds. The chosen response is accurate and appropriately uncertain; the rejected response is confident but wrong — fabricated statistics, fake citations, wrong facts, overclaimed certainty. Motivation Hallucination is the #1 reliability concern blocking enterprise LLM adoption. Models fail in predictable patterns: Inventing specific statistics with false… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-reduction-dpo-100k.texttext-generation100K<n<1M0 likes56 downloads2mo agoHugging Face29ssurface /hallucination-bert-spans Hallucination BERT Span Dataset Flat, one-row-per-span dataset intended for span/token-classification (BIO-tagging style) hallucination detection over agent tool-calling traces, derived from the same judging pipeline as the reasoning-distillation set in this collection. File ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join needed. Each row is one hallucinated span: span (verbatim text), type (taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.tabular10K<n<100K0 likes56 downloads2mo agoHugging Face30plm-hallubench /plm-hallubench PLM-HalluBench: A Multi-level Benchmark for Evaluating Hallucinations in Protein Language Models NeurIPS 2026 Evaluations and Datasets Track submission (double-blind). PLM-HalluBench is a benchmark for evaluating hallucination in protein language models (PLMs) — outputs that look like proteins but violate basic biophysics. It is organised around a three-level taxonomy that separates sequence-, structure-, and function-level failure modes, and it pairs a Factual track (BPHS against a… See the full description on the dataset page: https://huggingface.co/datasets/plm-hallubench/plm-hallubench.tabularother10K<n<100K0 likes52 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.