wmt24pp
Datasets
All datasets matching “wmt24pp”wmt24pp
WMT24++
This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in
the publication
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME.
If you are interested in the images of the source URLs for each document, please see here.
Schema
Each language pair is stored in its own jsonl file.
Each row… See the full description on the dataset page: https://huggingface.co/datasets/google/wmt24pp.wmt24pp-qtranslated
WMT24++ Translated by Quantized LLMs (wmt24pp-qtranslated)
This dataset accompanies the paper: [TBC]
It provides segment‑level translations and metadata for 55 languages × 110 directions produced by variants of the Llama 3.x and Qwen3 model families, each quantized with up to four post‑training quantization (PTQ) methods and two bit‑widths.
🌍 Dataset Structure
Split names : <src>-<tgt> (e.g. ar_EG-en, en-zu_ZA) ─ 110 in total
Columns :
• source_segment… See the full description on the dataset page: https://huggingface.co/datasets/bnjmnmarie/wmt24pp-qtranslated.wmt24pp
WMT24++
This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in
the publication
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME.
If you are interested in the images of the source URLs for each document, please see here.
Schema
Each language pair is stored in its own jsonl file.
Each row is… See the full description on the dataset page: https://huggingface.co/datasets/synquid/wmt24pp.wmt24pp
WMT24++
This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in
the publication
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME.
If you are interested in the images of the source URLs for each document, please see here.
Schema
Each language pair is stored in its own jsonl file.
Each row is… See the full description on the dataset page: https://huggingface.co/datasets/Lulu19971017/wmt24pp.wmt24pp-rm
WMT24++ Reference Translations for Romansh
Description
WMT24++ benchmark in Romansh (six varieties: Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader).
Paper: "Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader"
Code: https://github.com/ZurichNLP/romansh_mt_eval
Original WMT24++ benchmark (55 languages): https://huggingface.co/datasets/google/wmt24pp
The reference translation have been… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/wmt24pp-rm.wmt24pp-kr-reversed
WMT24++ Parallel Mix (Direction-Flipped)
Source dataset dir: /Users/inertia/Desktop/preprocessing/processed/wmt24pp_parallel_80_20
Transformation: direction flip per row (source/target swap)
Pair+sample_key overlap with source: 0
Primary(stats): {'en->ko': 450, 'ko->en': 449}
Auxiliary(flipped): ja->en, en->ja, zh->en, en->zh
Target primary ratio: 0.8000
Achieved primary ratio: 0.7998
Tag template: <{tgt_upper}>
Text template: {src_tag} {source} {tgt_tag} {target}… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/wmt24pp-kr-reversed.
