CoolFace
Datasetpublic

ML0037/ClosureBench

ClosureBench ClosureBench is a controlled benchmark for evaluating if LLMs respect explicit semantic contracts about missing information. It tests whether models distinguish absence-as-unknown, absence-as-false, and absence-as-false-only-in-complete-scopes under explicit open-world, closed-world, and locally closed-world contracts. The dataset includes the base benchmark and three extensions: Config full rows Description base 960 Main OWA/CWA/LCWA benchmark with… See the full description on the dataset page: https://huggingface.co/datasets/ML0037/ClosureBench.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes74downloads

ML0037/ClosureBench · main · files are served by the source, never re-hosted here