cai-bench
cai-semantic-equivalence-benchmark
Contradish CAI-Bench
The semantic equivalence benchmark from Contradish
Do AI systems give the same answer when the wording changes but the meaning does not?
Contradish CAI-Bench measures semantic invariance: whether an AI system remains behaviorally consistent across prompts that express the same intent in different words.
This Hugging Face release contains 420 human-readable prompt pairs across 19 domains. Contradish is the official benchmark runner, scoring… See the full description on the dataset page: https://huggingface.co/datasets/compressionawareintelligence/cai-semantic-equivalence-benchmark.cai-semantic-equivalence-benchmark
CAI Semantic Equivalence Benchmark
Version: 0.3
Pairs: 420
Domains: 19
License: MIT
A benchmark for measuring semantic invariance in language models. Tests whether a model gives the same answer when the same question is rephrased.
This is the evaluation dataset behind the CAI Semantic Equivalence Benchmark and scored by contradish using CAI Strain v2.
What it tests
Most LLM benchmarks test accuracy. This one tests consistency. A model passes when it gives… See the full description on the dataset page: https://huggingface.co/datasets/theworkforceof/cai-semantic-equivalence-benchmark.
