datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jev-calibration-statistics
Confidence statistics for Jev and a self-judging Gemma 4 E2B
Aggregate statistics on the confidence scores from two judges in a retrieval benchmark: TypeSafe's Jev, pinned to jev-1.13.0, and Gemma 4 E2B judging its own work.
End to end, the pipeline with Jev making every decision did not beat the same pipeline with no judge: it scored 0.612 against 0.740, missed its main pre-registered bar, made about the same number of mistakes on questions both answered, and lost because it… See the full description on the dataset page: https://huggingface.co/datasets/clduab11/jev-calibration-statistics.Calibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
urban_air_quality_satellite_calibrationCalibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
lambda-calibration-outputs
