CoolFace
Datasetpublic

kilizi/FactGuard

FactGuard-Bench FactGuard-Bench is a bilingual long-context benchmark for evaluating and improving whether language models answer only when the supplied document contains sufficient evidence. It contains English and Chinese examples from the book and legal domains, with contexts extending to approximately 128K in the legacy character-based construction buckets. The benchmark accompanies: Towards Reliable Long-Context Reasoning: Detecting Unanswerable Questions via FactGuard… See the full description on the dataset page: https://huggingface.co/datasets/kilizi/FactGuard.

sourceHugging Facecc-by-4.0updated 25d agoView on Hugging Face
0likes95downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

kilizi/FactGuard · main · files are served by the source, never re-hosted here