klusai/open-weight-judges-eval
Open-Weight Judges Evaluation Dataset Evaluation outputs from a study on cross-task transferability of family-diverse open-weight judge panels for synthetic text evaluation. Overview This dataset contains ~6,180 judge evaluations across three synthetic-text tasks, produced by a three-member open-weight panel, proprietary baselines (GPT-o4-mini, Gemini 2.5 Flash), a specialized open evaluator (Atla Selene Mini), and four bias audits. Tasks Task… See the full description on the dataset page: https://huggingface.co/datasets/klusai/open-weight-judges-eval.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face