esperanto
esperanto-mt-parallel-v13
esperanto-mt-parallel v13
EN<->EO parallel training corpus. 5,025,333 rows after dedup.
What changed vs v12 (jensjepsen/esperanto-mt-parallel)
Dropped Helsinki-NLP/opus-100 en-eo train split (144,549 rows).
It is an aggregated multilingual blob that bundles KDE4/GNOME/Ubuntu
.po localization pairs without src labels. In v12 this caused a
systematic MT failure mode: capitalized-fragment-no-terminal-punct
inputs collapsed to memorized UI labels
(e.g. "@ info:… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-mt-parallel-v13.esperanto-hpltesperanto-mt-parallel
ccmatrix_filtered_dedup
xlent_dedup
tatoeba_train_dedup
wikimatrix_dedup
opus100_train_dedup
bible_uedin_dedup
opensubtitles_v2024_dedup
wikimedia_dedup
ted2020_dedup
opusbooks_train_dedup
HuggingFaceFW/finetranslations (epo_Latn)
jensjepsen/esperanto-sciq
jensjepsen/esperanto-balanced-copa
jensjepsen/esperanto-piqa
jensjepsen/esperanto-mmlu
jensjepsen/esperanto-triviaqa
esperanto-sciqesperanto-metamath-gsmesperanto-boolq-questions
esperanto-boolq-questions
BoolQ questions (train + validation, 12,697 rows) translated from
English to Esperanto by
jensjepsen/eo-mt-v13-large-bidir,
with round-trip quality metadata for filtering.
Row schema
field
description
orig_idx
original BoolQ row index (train first, then validation)
split
source split (train / validation)
en_orig
raw BoolQ question (lowercase, no ?, as in google/boolq)
en_preproc
preprocessed input fed to MT: spaCy… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-boolq-questions.
