CoolFace
Datasetpublic

kilizi/FactGuard

FactGuard-Bench FactGuard-Bench is a bilingual long-context benchmark for evaluating and improving whether language models answer only when the supplied document contains sufficient evidence. It contains English and Chinese examples from the book and legal domains, with contexts extending to approximately 128K in the legacy character-based construction buckets. The benchmark accompanies: Towards Reliable Long-Context Reasoning: Detecting Unanswerable Questions via FactGuard… See the full description on the dataset page: https://huggingface.co/datasets/kilizi/FactGuard.

sourceHugging Facecc-by-4.0updated 24d agoView on Hugging Face
0likes95downloads
5 commits on main
0f0c69024d ago

Fix dataset metadata for load_dataset

kilizi
e49360e24d ago

Update dataset loading examples and single-file split paths

kilizi
c0c0e6724d ago

Consolidate each dataset split into one Parquet file

kilizi
2f922a724d ago

Publish FactGuard-Bench with train, validation, and test splits

kilizi
39e435d24d ago

initial commit

kilizi