datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esperanto-orca-math-full
Esperanto Orca-Math (full, unfiltered)
197,849 EO chain-of-thought math reasoning examples translated from
microsoft/orca-math-word-problems-200k via jensjepsen/eo-mt-v5b
(MarianMT), with v5b operator-drop artifacts repaired post-hoc.
This is the unfiltered companion to jensjepsen/esperanto-orca-math
(which is filtered to Q+A ≤ 512 morpheme-tokens for our small student).
Use this one if your model has a wider context window — it includes
the long reasoning chains that didn't fit… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-orca-math-full.esperanto-alpaca-distill
esperanto-alpaca-distill
Esperanto Alpaca distill: 98,329 instruction/response pairs from yahma/alpaca-cleaned, distilled through LiquidAI/LFM2.5-350M (EN answers), then translated EN→EO via jensjepsen/eo-mt-v5b. Format: messages.
esperanto-orca-math
Esperanto Orca-Math
Chain-of-thought math reasoning examples in Esperanto, translated from
microsoft/orca-math-word-problems-200k via jensjepsen/eo-mt-v5b
(MarianMT en→eo). v5b's occasional operator-drop artifacts
(e.g. 6 * 9 → 6 9) were repaired post-hoc by detecting A B = C
patterns where A * B == C and re-inserting the *.
Format: {"messages": [{"role": "user", "content": Q_eo}, {"role": "assistant", "content": A_eo}]}
Files
train.jsonl (128,689 rows) —… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-orca-math.
