datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esperanto-mt-parallel-v13
esperanto-mt-parallel v13
EN<->EO parallel training corpus. 5,025,333 rows after dedup.
What changed vs v12 (jensjepsen/esperanto-mt-parallel)
Dropped Helsinki-NLP/opus-100 en-eo train split (144,549 rows).
It is an aggregated multilingual blob that bundles KDE4/GNOME/Ubuntu
.po localization pairs without src labels. In v12 this caused a
systematic MT failure mode: capitalized-fragment-no-terminal-punct
inputs collapsed to memorized UI labels
(e.g. "@ info:… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-mt-parallel-v13.esperanto-mt-parallel
ccmatrix_filtered_dedup
xlent_dedup
tatoeba_train_dedup
wikimatrix_dedup
opus100_train_dedup
bible_uedin_dedup
opensubtitles_v2024_dedup
wikimedia_dedup
ted2020_dedup
opusbooks_train_dedup
HuggingFaceFW/finetranslations (epo_Latn)
jensjepsen/esperanto-sciq
jensjepsen/esperanto-balanced-copa
jensjepsen/esperanto-piqa
jensjepsen/esperanto-mmlu
jensjepsen/esperanto-triviaqa
esperanto-hpltesperanto-sciqesperanto-metamath-gsmesperanto-boolq-questions
esperanto-boolq-questions
BoolQ questions (train + validation, 12,697 rows) translated from
English to Esperanto by
jensjepsen/eo-mt-v13-large-bidir,
with round-trip quality metadata for filtering.
Row schema
field
description
orig_idx
original BoolQ row index (train first, then validation)
split
source split (train / validation)
en_orig
raw BoolQ question (lowercase, no ?, as in google/boolq)
en_preproc
preprocessed input fed to MT: spaCy… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-boolq-questions.esperanto-word-problems-v3esperanto-gsm8kesperanto-arcesperanto-balanced-copaalpaca_esperanto_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_esperanto_taco.esperanto-sft-atomic-iclesperanto-squadEsperantoBenchThis is a benchmark for testing knowledge of the Esperanto language. It contains 651 questions regarding Esperanto vocabulary and grammar. All questions are in 4-answer multiple choice format.
The questions and answers are all based on the book "A Complete Grammar of Esperanto" by Ivy Kellerman Reed. Source: https://www.gutenberg.org/ebooks/7787
As of publishing this on June 23rd, 2025, it seems that this benchmark is already saturated! Neat.
gpt-4.1-nano: 83.87%
gpt-4.1-mini: 93.39%
gpt-4.1:… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/EsperantoBench.esperanto-sft-wikidata-iclesperanto
Dataset Card for "esperanto"
More Information needed
esperanto-arithmetic-cotesperanto-piqaesperanto-word-problems-v2esperanto-sentencesesperanto-sft-atomic-qaesperanto_pt_nondedupesperanto-sft-morphology-iclesperanto-orca-math-full
Esperanto Orca-Math (full, unfiltered)
197,849 EO chain-of-thought math reasoning examples translated from
microsoft/orca-math-word-problems-200k via jensjepsen/eo-mt-v5b
(MarianMT), with v5b operator-drop artifacts repaired post-hoc.
This is the unfiltered companion to jensjepsen/esperanto-orca-math
(which is filtered to Q+A ≤ 512 morpheme-tokens for our small student).
Use this one if your model has a wider context window — it includes
the long reasoning chains that didn't fit… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-orca-math-full.esperanto-word-problemsesperanto_bible_eo
Esperanto Bible (Londona Biblio)
Description
The Londona Biblio (London Bible) is a complete translation of the Bible into Esperanto, the international auxiliary language created by L. L. Zamenhof. The translation was prepared by a team of Esperanto-speaking scholars and published by the British and Foreign Bible Society. It includes the Protestant canon (66 books). This translation is significant as it makes the Bible accessible to Esperanto speakers worldwide… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/esperanto_bible_eo.esperanto-sft-dollyesperanto-alpaca-distill
esperanto-alpaca-distill
Esperanto Alpaca distill: 98,329 instruction/response pairs from yahma/alpaca-cleaned, distilled through LiquidAI/LFM2.5-350M (EN answers), then translated EN→EO via jensjepsen/eo-mt-v5b. Format: messages.
alpaca-esperanto-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-esperanto-cleaned.esperanto-gutenberg
