CoolFace
Datasetpublic

s-emanuilov/rivers-qa

Dataset for Evaluating Hallucination Mitigation in Language Models This dataset contains structured question-answer pairs derived from U.S. river knowledge, designed to test factual grounding and epistemic discipline in LLMs. The dataset includes raw data for 9,538 river entities with 21 attributes (length, discharge, source/mouth locations, etc.) and an evaluation set of 17,726 question-answer pairs targeting factual relationships. Citation… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-qa.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes10downloads
Dataset Card

Dataset for Evaluating Hallucination Mitigation in Language Models

This dataset contains structured question-answer pairs derived from U.S. river knowledge, designed to test factual grounding and epistemic discipline in LLMs. The dataset includes raw data for 9,538 river entities with 21 attributes (length, discharge, source/mouth locations, etc.) and an evaluation set of 17,726 question-answer pairs targeting factual relationships.

Citation

bibtex
@article{ackermann2025stemming,
  title={Stemming Hallucination in Language Models Using a Licensing Oracle},
  author={Ackermann, Richard and Emanuilov, Simeon},
  year={2025}
}