jev
Datasets
All datasets matching “jev”jev-bench
jev-bench
Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist.
22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · v0.1.1
Repo & engine · Source rationale · What we verified about Jev's API · Other independent Jev evaluations
jev-1.13.0 on every test record: crisp, grounded decisions land in the accurate-and-calibrated corner; ordinal ratings and anything humans disagree about do not.… See the full description on the dataset page: https://huggingface.co/datasets/Praveenrajus/jev-bench.jev-stage2-image-beans-pilot
Beans: one natural question per image
Open the corrected preview.
natural_v4 is the recommended and default preview: 100 original images, 100 rows, one three-way condition-class Choice question per image. All targets come directly from the source labels column (34 angular leaf spot, 33 bean rust, 33 healthy). Original image bytes and source annotations are unchanged.
Example question: “Which source-defined condition class describes the bean leaf?” Options: angular_leaf_spot… See the full description on the dataset page: https://huggingface.co/datasets/FaroukMoc2/jev-stage2-image-beans-pilot.jev-luna-pagerduty-trigger
Jev vs Luna as a PagerDuty trigger
Synthetic checkout/payments log stream with gold labels from PagerDuty alerting principles: page only if a human must act now. TypeSafe’s Jev (typesafe-ai/jev) and GPT-5.6 Luna (openai/gpt-5.6-luna) both ran on Vercel AI Gateway. There is no ERROR auto-page.
This is not production traffic and not the Loghub junk-filter benchmark.
Write-up: https://github.com/reachjalil/jevlogs/blob/hf-benchmark/docs/article/jev-vs-luna-pagerduty.md
Code:… See the full description on the dataset page: https://huggingface.co/datasets/reachjalil/jev-luna-pagerduty-trigger.jevlogs-log-triage-benchmark
Jev Logs log-triage benchmark
A labeled evaluation of Jev Logs on sanitized public logs. Jev Logs asks TypeSafe’s Jev, through Vercel AI Gateway, whether a log line is worth sending to an expensive reasoning model. This dataset is a public, token-accounted measurement of that routing decision, including the 0.3.0 in-memory cache and local retain rules.
This is not a production-log study. Labels come from Loghub. HDFS labels are block-level, then joined onto every line that… See the full description on the dataset page: https://huggingface.co/datasets/reachjalil/jevlogs-log-triage-benchmark.Open-Jev
Open-Jev: typed decision datasets
Open-Jev turns a state and a question into a typed decision: a yes/no probability, a distribution over choices, independent label probabilities, or a discrete numeric/ordinal decision. This repository publishes twelve separate, frozen data configs from the Open-Jev project, together with original manifests, exact raw records, source code and reconstruction instructions.
These are controlled, mostly synthetic tasks and reference labels. They are… See the full description on the dataset page: https://huggingface.co/datasets/ZefanCai/Open-Jev.open-jev-laya-bench
