CoolFace
Datasetpublic

benchmark-data/hallucination-traps

Hallucination Traps A curated benchmark dataset consisting of intentionally misleading prompts designed to evaluate hallucination behavior in language models. Each prompt appears plausible at first glance but contains a subtle false premise, nonexistent entity, or incorrect factual assumption. The expected behavior is that the model should either refuse, express uncertainty, or explicitly identify the incorrect premise rather than hallucinate a confident but false answer.… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-data/hallucination-traps.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes15downloads
Dataset Card

Hallucination Traps

A curated benchmark dataset consisting of intentionally misleading prompts designed to evaluate hallucination behavior in language models.

Each prompt appears plausible at first glance but contains a subtle false premise, nonexistent entity, or incorrect factual assumption. The expected behavior is that the model should either refuse, express uncertainty, or explicitly identify the incorrect premise rather than hallucinate a confident but false answer.

Intended Use

This dataset is intended for:

  • —Evaluating factual robustness
  • —Testing hallucination resistance
  • —Comparing refusal and uncertainty behaviors across models
  • —Benchmarking model improvements over time

Structure

Each entry contains:

  • —A prompt containing a false or misleading premise
  • —An expected behavior description
  • —A category label describing the hallucination type

License

This dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

Citation

If you use this dataset in academic work or benchmarks, please cite or link to:

https://huggingface.co/datasets/benchmark-data/hallucination-traps