YangL1122/DetailBench
DetailBench A Benchmark for Detail Hallucination in Long Regulatory Documents DetailBench is a benchmark for evaluating and mitigating detail hallucination in LLM outputs on long regulatory documents. Overview Large language models frequently produce detail hallucinations—subtle errors in threshold values, units, scopes, obligation levels, and conditions—when processing long regulatory documents. DetailBench provides: 322 source documents (172 real + 150… See the full description on the dataset page: https://huggingface.co/datasets/YangL1122/DetailBench.
This repository belongs to YangL1122 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
