CoolFace
Datasetpublic

kernel-14/SemanticAlign-Bench

SemanticAlign-Bench A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering. The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes573downloads
paper.pdf4 linesDownload Raw Back to sam2
1version https://git-lfs.github.com/spec/v12oid sha256:6c07c6ecf2f592e3f7f9833409b8ec63232e2193b4ab4edb7db575805f27c0283size 119982844