mt-evaluation
romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.en_be_mt_datasets_evaluation
Overview
This is a small dataset of English-Belarusian sentence pairs sampled from the largest parallel corpora in OPUS (100 random instances from each of the following: NLLB, HPLT, CCMatrix, CCAligned) and manually labeled for correctness by a speaker of Belarusian. The taxonomy of labels follows Kreutzer et al. 2022:
CC: correct translation, natural sentence
CB: correct translation, boilerplate or low quality
CS: correct translation, short
X: incorrect translation
WL: wrong… See the full description on the dataset page: https://huggingface.co/datasets/somerandomguyontheweb/en_be_mt_datasets_evaluation.mt-human-evaluation-da
