datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PRE-HAL
PRE-HAL: Multimodal Hallucination Evaluation Benchmark
Dataset Summary
PRE-HAL is a visual question answering (VQA) dataset designed to evaluate and mitigate hallucination in Multimodal Large Language Models (MLLMs). It focuses on testing the model's ability to distinguish between visual perception and parametric knowledge, specifically targeting various hallucination types.
Data Instances
Each instance represents a multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/TerryHWong/PRE-HAL.HALP-Bench
HALP-Bench
HALP-Bench is the evaluation benchmark released with the paper
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
(EACL 2026). It aggregates 10,000 image-question pairs drawn from six public sources into a
single, uniformly-formatted set with stable image IDs.
📄 Paper (ACL Anthology): https://aclanthology.org/2026.eacl-long.287/
📄 Preprint (arXiv): https://arxiv.org/abs/2603.05465
💻 Code: https://github.com/Zesearch/HALP… See the full description on the dataset page: https://huggingface.co/datasets/Zesearch/HALP-Bench.HaloQuestThis Dataset introduces challenging pictures for assessing MLLM hallucinations. It includes questions and ground-truth answers.
Original Paper: (https://arxiv.org/abs/2407.15680)[https://arxiv.org/abs/2407.15680]
Code Repo: (https://github.com/google/haloquest/)[https://github.com/google/haloquest/]
Language-Vision-Hallucinations
Dataset for Techen Project 095280
A comprehensive dataset for the Techen Project, focused on examining hallucinations in multi-modal AI-generated text by investigating model uncertainty, text generation patterns, and linguistic factors.
Columns Overview
image_link: URL to the image associated with each data row.
temperature: Temperature setting for text generation, controlling output randomness.
description: Text generated by the model for each image, using the… See the full description on the dataset page: https://huggingface.co/datasets/wrom/Language-Vision-Hallucinations.JourneyBench_Hallucinationhalal-cert
