datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
es_eu-flaskBenchmark used in our paper, "Towards Reliable Multilingual Judge Models: An Empirical Study."
The full code is available in the GitHub repository hitz-zentroa/mJudge.
The original English partition from which this benchmark was derived can be found in FLASK.
Dataset Variants
flask_es: All fields translated into Spanish.
flask_eu: All fields translated into Basque.
flask_io_es: Only the model input and output to be evaluated are translated into Spanish; all other fields remain in… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/es_eu-flask.kaistai-flask
FLASK: Fine-grained Language Model Evaluation Based on Alignment Skill Sets
FLASK evaluation dataset. Originally published here.
Changes Made to the dataset
Some changes were made to the dataset to make it compatible for upload to Huggingface.
-1s were replaced with "Unknown" in their respective columns
Columns that appeared inconsistent with the majority of the data format were removed
mmlu
subcategory
subdomain
deterministic
mcqa
subtag
skill_domain
skill_subdomain
