datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wmt24pp
WMT24++
This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in
the publication
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME.
If you are interested in the images of the source URLs for each document, please see here.
Schema
Each language pair is stored in its own jsonl file.
Each row… See the full description on the dataset page: https://huggingface.co/datasets/google/wmt24pp.wmt24pp
WMT24++
This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in
the publication
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME.
If you are interested in the images of the source URLs for each document, please see here.
Schema
Each language pair is stored in its own jsonl file.
Each row is… See the full description on the dataset page: https://huggingface.co/datasets/Lulu19971017/wmt24pp.wmt24pp
WMT24++
This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in
the publication
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME.
If you are interested in the images of the source URLs for each document, please see here.
Schema
Each language pair is stored in its own jsonl file.
Each row is… See the full description on the dataset page: https://huggingface.co/datasets/synquid/wmt24pp.wmt24pp-kr-reversed
WMT24++ Parallel Mix (Direction-Flipped)
Source dataset dir: /Users/inertia/Desktop/preprocessing/processed/wmt24pp_parallel_80_20
Transformation: direction flip per row (source/target swap)
Pair+sample_key overlap with source: 0
Primary(stats): {'en->ko': 450, 'ko->en': 449}
Auxiliary(flipped): ja->en, en->ja, zh->en, en->zh
Target primary ratio: 0.8000
Achieved primary ratio: 0.7998
Tag template: <{tgt_upper}>
Text template: {src_tag} {source} {tgt_tag} {target}… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/wmt24pp-kr-reversed.wmt24pp_xx_to_pt_10k
WMT24PP XX to PT 10k
This is a dataset of translations from various source languages into Portuguese. The pairs are derived from google/wmt24pp, where the same 998 English sentences are each translated into many target languages. For every sentence, we take its translation in another language as the source and its Portuguese translation as the target, yielding {source language}→Portuguese pairs across many languages. Note that both sides are translations of the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wmt24pp_xx_to_pt_10k.wmt24pp-kr
WMT24++ Parallel Mix
Primary(all): en->ko, ko->en
Primary mode: disjoint_halves
Auxiliary(sampled, no oversampling): en->ja, ja->en, en->zh, zh->en
Target primary ratio: 0.8000
Achieved primary ratio: 0.7998
Tag template: <{tgt_upper}>
Text template: {src_tag} {source} {tgt_tag} {target}
Normalize doubled quotes: True
Strip control chars: True
Drop bad source: True
Drop canary: True
Drop @user handles: True
Drop one-word sentences: True
Columns
id, dataset, pair_config… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/wmt24pp-kr.wmt24pp-ce
WMT24++ Reference Translations for Chechen
Description
WMT24++ benchmark in Chechen
Original WMT24++ benchmark (55 languages): https://huggingface.co/datasets/google/wmt24pp
The reference translation have been created by human translator based on the Russian version of WMT24++
The dataset uses both cases of Cyrillic Palochka Letter where it is grammatically correct.
For preparation of the Chechen version of the dataset we hired a professional native speaker… See the full description on the dataset page: https://huggingface.co/datasets/NM-development/wmt24pp-ce.wmt24pp-en-bn
WMT24++ English-Bengali Filtered Subset
This dataset is a filtered subset of google/wmt24pp using en-bn_IN.jsonl as the source file.
Filtering rules:
keep rows where is_bad_source is false
keep rows with non-empty source
keep rows with non-empty target
Result:
input rows: 998
kept rows: 960
removed rows: 38
Columns are preserved from the source JSONL.
