CoolFace
Datasetpublic

anote-ai/IVR-pilot-benchmark

IntentSpec Benchmark — Data Supplement This archive contains the benchmark data used to compute Intent Violation Rate (IVR) in the paper: 49 tasks, each derived from a HumanEval problem and extended with an ambiguous/gold prompt pair and a decomposed set of executable constraints. Files spec_pairs.jsonl The benchmark itself — one JSON object per line, one line per task. This is the file consumed directly by the evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/IVR-pilot-benchmark.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes26downloads

anote-ai/IVR-pilot-benchmark · main · files are served by the source, never re-hosted here