Plumloom/evaluation-reliability-benchmark
Plumloom Evaluation Reliability Benchmark Public results from Plumloom’s research on reliability in single-turn AI chat evaluations. This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs. Private prompts, rubrics, model responses, and production execution details are intentionally excluded.
Plumloom Evaluation Reliability Benchmark
Public results from Plumloom’s research on reliability in single-turn AI chat evaluations.
This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs.
Private prompts, rubrics, model responses, and production execution details are intentionally excluded.
