datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.wmt-mqmThis dataset contains all MQM human annotations from WMT Metrics Shared Tasks from 2020 to 2024.
The data is organized into different multiple columns, all should contain the following columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
This dataset is taken directly from the original WMT repository.
wmt24-mqm-qewmt24-mqm-qeTurkDoc-MT-MQM-Annotations
