eval
Datasets
All datasets matching “eval”evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}mbppplusVideo-MMEmmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.EEE_datastore
Every Eval Ever Datastore
A community database of AI evaluation results, all in one schema. Scores scraped from
leaderboards, pulled out of papers, and produced by local evaluation runs are stored in a
single record format, so results from different sources can be compared, joined, and reused
instead of re-scraped. This dataset is the data itself: one JSON record per
model per evaluation run — which may carry several scored results — with optional
per-sample companion files.… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/EEE_datastore.humanevalplus

