datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies).
It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains.
Watch a brief 4 minutes-long video.
Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.wmt-mqm-error-spans
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context in a form of error spans. Moreover, it contains some hallucinations used in the training of XCOMET models.
Please note that this is not an official release of the data and the original data can be found here.
The data is organised into 8 columns:
src: input text
mt: translation
ref: reference translation
annotations: List… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-error-spans.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.peer_qt21-de-en-mqm
QT21 De-En MQM Task from the PEER Benchmark (Performance Evaluation of Edit Representations)
Description from the benchmark paper:
A subset of 1,800 examples of the De-En QT21 dataset, annotated with details about the edits performed, namely the reason why each edit was applied. Since the dataset contains a large number of edit labels, we select the classes that are present in at least 100 examples and generate a modified version of the dataset for our purposes. Examples where no… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/peer_qt21-de-en-mqm.mqm-translation-gold
Alconost MQM Translation Quality Dataset
A growing collection of professional MQM (Multidimensional Quality Metrics) annotations for machine translation evaluation.
Dataset Description
This dataset contains human expert annotations of machine translation outputs using the MQM framework - the same methodology used in WMT (Workshop on Machine Translation) human evaluation campaigns.
All annotations are performed by trained linguists with native/near-native proficiency. Data… See the full description on the dataset page: https://huggingface.co/datasets/alconost/mqm-translation-gold.wmt24-mqmwmt-mqmThis dataset contains all MQM human annotations from WMT Metrics Shared Tasks from 2020 to 2024.
The data is organized into different multiple columns, all should contain the following columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
This dataset is taken directly from the original WMT repository.
GPT-4_FO-EN_parallel_blog_sentences_MQMThis is dataset contains 425 Faroese-to-English parallel sentences generated by GPT-4 that have been annotated by a single native speaker of Faroese using the Multidimensional Quality Metrics framework (MQM). The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences and human evaluation, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences_MQM.wmt24-mqm-qewmt24-mqmGPT-4_FO-EN_parallel_news_sentences_MQMThis is dataset contains 425 Faroese-to-English parallel sentences generated by GPT-4 that have been annotated by a single native speaker of Faroese using the Multidimensional Quality Metrics framework (MQM). The Faroese text is news text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences and human evaluation, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_news_sentences_MQM.wmt24-mqm-qeTurkDoc-MT-MQM-AnnotationsWMT23_MQM_PairwiseWMT22_MQM_Pairwisetoy_wmt24_mqm
