mqm
sentinel-cand-mqmsentinel-src-mqmsentinel-ref-mqmmqmsvaa8isfhvq7nuppqccvio3o1_6c30f82d-5b6f-4f5f-8901-6a6f5a5153bcmqmsvaa8isfhvq7nuppqccvio3o1_6bb8aa47-36a3-4515-9e90-67e8e111c8c0mqmsvaa8isfhvq7nuppqccvio3o1_22739782-5f5d-4154-a302-ace25fca24d2comet-bio-mqm-n6jrlwttmqmsvaa8isfhvq7nuppqccvio3o1_cea7b6fc-5d7d-400a-a98e-d67abf603e4d
Datasets
All datasets matching “mqm”bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies).
It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains.
Watch a brief 4 minutes-long video.
Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.wmt-mqm-error-spans
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context in a form of error spans. Moreover, it contains some hallucinations used in the training of XCOMET models.
Please note that this is not an official release of the data and the original data can be found here.
The data is organised into 8 columns:
src: input text
mt: translation
ref: reference translation
annotations: List… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-error-spans.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.peer_qt21-de-en-mqm
QT21 De-En MQM Task from the PEER Benchmark (Performance Evaluation of Edit Representations)
Description from the benchmark paper:
A subset of 1,800 examples of the De-En QT21 dataset, annotated with details about the edits performed, namely the reason why each edit was applied. Since the dataset contains a large number of edit labels, we select the classes that are present in at least 100 examples and generate a modified version of the dataset for our purposes. Examples where no… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/peer_qt21-de-en-mqm.mqm-translation-gold
Alconost MQM Translation Quality Dataset
A growing collection of professional MQM (Multidimensional Quality Metrics) annotations for machine translation evaluation.
Dataset Description
This dataset contains human expert annotations of machine translation outputs using the MQM framework - the same methodology used in WMT (Workshop on Machine Translation) human evaluation campaigns.
All annotations are performed by trained linguists with native/near-native proficiency. Data… See the full description on the dataset page: https://huggingface.co/datasets/alconost/mqm-translation-gold.wmt24-mqm
