CoolFace
Datasetpublic

ymoslem/wmt-da-human-evaluation-long-context

Dataset Summary Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
8likes398downloads
Dataset Card

Dataset Summary

Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the augmentation can be found here.

This dataset contains all DA human annotations from previous WMT News Translation shared tasks. It extends the sentence-level dataset RicardoRei/wmt-da-human-evaluation, split into train and test. Moreover, the raw column is normalized to be between 0 and 1 using this function.

The data is organised into 8 columns:

  • lp: language pair
  • src: input text
  • mt: translation
  • ref: reference translation
  • raw: direct assessment
  • domain: domain of the input text (e.g. news)
  • year: collection year
  • sents: number of sentences in the text

You can also find the original data for each year in the results section: https://www.statmt.org/wmt{YEAR}/results.html e.g: for 2020 data: https://www.statmt.org/wmt20/results.html

Python usage:

python
from datasets import load_dataset
dataset = load_dataset("ymoslem/wmt-da-human-evaluation-long-context")

There is no standard train/test split for this dataset, but you can easily split it according to year, language pair or domain. e.g.:

python
# split by year
data = dataset.filter(lambda example: example["year"] == 2022)

# split by LP
data = dataset.filter(lambda example: example["lp"] == "en-de")

# split by domain
data = dataset.filter(lambda example: example["domain"] == "news")

Note that most data is from the News domain.

Citation Information

If you use this data please cite the WMT findings from previous years: