kilizi/FactGuard
FactGuard-Bench FactGuard-Bench is a bilingual long-context benchmark for evaluating and improving whether language models answer only when the supplied document contains sufficient evidence. It contains English and Chinese examples from the book and legal domains, with contexts extending to approximately 128K in the legacy character-based construction buckets. The benchmark accompanies: Towards Reliable Long-Context Reasoning: Detecting Unanswerable Questions via FactGuard… See the full description on the dataset page: https://huggingface.co/datasets/kilizi/FactGuard.
Fix dataset metadata for load_dataset
Update dataset loading examples and single-file split paths
Consolidate each dataset split into one Parquet file
Publish FactGuard-Bench with train, validation, and test splits
initial commit
