CoolFace
Datasetpublic

YangL1122/DetailBench

DetailBench A Benchmark for Detail Hallucination in Long Regulatory Documents DetailBench is a benchmark for evaluating and mitigating detail hallucination in LLM outputs on long regulatory documents. Overview Large language models frequently produce detail hallucinations—subtle errors in threshold values, units, scopes, obligation levels, and conditions—when processing long regulatory documents. DetailBench provides: 322 source documents (172 real + 150… See the full description on the dataset page: https://huggingface.co/datasets/YangL1122/DetailBench.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
0likes55downloads

YangL1122/DetailBench · main · files are served by the source, never re-hosted here