datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdsuite-delphi-result
GDsuite results — Delphi model collection
GDsuite evaluation results for
the marin-community/delphi
model collection.
Contents
summary.jsonl — tidy per-task metrics (14256 rows). One row per
(model, family, task, metric):
metric = hard_acc (5 logprob families) — fraction of items where
P(correct) > P(incorrect); higher ⇒ resists the misleading pattern.
metric = correct_log_prob (5 logprob families) — mean
teacher-forced log probability of the correct answer.
metric =… See the full description on the dataset page: https://huggingface.co/datasets/WillHeld/gdsuite-delphi-result.gdsc_llm
