datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.MuST-C-and-WMT16-de-enwmt-sqm-human-evaluation
Dataset Summary
In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA).
The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.wmt24-mqmwmt-mqmThis dataset contains all MQM human annotations from WMT Metrics Shared Tasks from 2020 to 2024.
The data is organized into different multiple columns, all should contain the following columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
This dataset is taken directly from the original WMT repository.
wmt24-mqm-qewmt24-mqmwmt22-biomedwmt24-mqm-qewmt-da-en-tr_tr-enwmt23-single-sentences-en-deThis is the test data from the WMT 23 general MT task (news-systems), but segmented into single sentences. It includes references to the original text segment ids.
This dataset only includes en-de and de-en data, with around 2k single sentence examples per direction.
The original data is taken from this GitHub repo.
This is the publication about the WMT general MT findings:
@inproceedings{wmt-2023-findings,
title = "Findings of the 2023 Conference on Machine Translation ({WMT}23): {LLM}s… See the full description on the dataset page: https://huggingface.co/datasets/yawnick/wmt23-single-sentences-en-de.WMT_EXT_DATAtoy_wmt24_mqmwmt_pluswmt-data-en-frwmt16_biomed_testwmt16_biomed_gold
