tedbelford/cliffs-and-slopes-semantic-cache
Cliffs and Slopes — Semantic-Cache Benchmark Benchmark for the paper "Cliffs and Slopes: Why Semantic Caches Fail Structurally, Not Just Through Miscalibration" (under submission to Information and Software Technology). A semantic cache answers a new query with the stored answer of a "similar enough" past query. This benchmark measures where — and why — threshold-on-similarity caching fails structurally: the task requires an equivalence relation over queries, but is implemented… See the full description on the dataset page: https://huggingface.co/datasets/tedbelford/cliffs-and-slopes-semantic-cache.
Sync human_gold.json analysis_rules wording (analysis set = three definite labels)
Human audit v2 analysis rules: exclude 3 double-UNSURE (n=1497), cluster-bootstrap CIs by seed; audited (not validated) wording + grounding scope note; card/datasheet/stats updated
Add human gold-subset validation: 3-annotator blind labels, protocol, guide, agreement stats (Fleiss k=0.76, 98.5% label agreement); update card + datasheet
Initial release: 46,214-pair benchmark, taxonomy v1.0, prompts, figures, datasheet
initial commit
