dyfan/halluhard-seed-questions
This is the seed question set for the benchmark HalluHard We design some difficult seed questions that elicit a citation-grounded multi-turn open-ended generation, covering four domains: legal, research, medical, and coding. Below is the Top15 models on our leaderboard. Please check our website for more models and turn-wise/domain-wise statistics! Rank Model Hallucination Rate Legal Research Medical Coding 1 Claude-Opus-4.5-Web-Search 30.2 33.0 29.6 29.2 29.0 2… See the full description on the dataset page: https://huggingface.co/datasets/dyfan/halluhard-seed-questions.
This is the seed question set for the benchmark HalluHard
We design some difficult seed questions that elicit a citation-grounded multi-turn open-ended generation, covering four domains: legal, research, medical, and coding.
Below is the Top15 models on our leaderboard. Please check our website for more models and turn-wise/domain-wise statistics!
Paper: https://arxiv.org/abs/2602.01031
Code: https://github.com/epfml/halluhard
If you use HalluHard in your research, please cite:
@misc{fan2026halluhardhardmultiturnhallucination,
title={HalluHard: A Hard Multi-Turn Hallucination Benchmark},
author={Dongyang Fan and Sebastien Delsad and Nicolas Flammarion and Maksym Andriushchenko},
year={2026},
eprint={2602.01031},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2602.01031},
}