jevbench
Datasets
All datasets matching “jevbench”jev-bench
jev-bench
Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist.
22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · v0.1.1
Repo & engine · Source rationale · What we verified about Jev's API · Other independent Jev evaluations
jev-1.13.0 on every test record: crisp, grounded decisions land in the accurate-and-calibrated corner; ordinal ratings and anything humans disagree about do not.… See the full description on the dataset page: https://huggingface.co/datasets/Praveenrajus/jev-bench.jev-bench
jev-bench
A small multiple-choice set for measuring a model that returns the probability of each option
instead of writing an answer — the Jev / TypeSafe System One style of API, where a request carries
a state and a question with named options and the response carries a distribution over them.
Ordinary multiple-choice benchmarks score the text a model generates. That says nothing about
whether the probability attached to the answer means anything, which is the whole point of… See the full description on the dataset page: https://huggingface.co/datasets/kishida/jev-bench.
