CoolFace
Datasetpublic

Plumloom/evaluation-reliability-benchmark

Plumloom Evaluation Reliability Benchmark Public results from Plumloom’s research on reliability in single-turn AI chat evaluations. This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs. Private prompts, rubrics, model responses, and production execution details are intentionally excluded.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes46downloads
Dataset Card

Plumloom Evaluation Reliability Benchmark

Public results from Plumloom’s research on reliability in single-turn AI chat evaluations.

This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs.

Private prompts, rubrics, model responses, and production execution details are intentionally excluded.