datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esperanto-boolq-questions
esperanto-boolq-questions
BoolQ questions (train + validation, 12,697 rows) translated from
English to Esperanto by
jensjepsen/eo-mt-v13-large-bidir,
with round-trip quality metadata for filtering.
Row schema
field
description
orig_idx
original BoolQ row index (train first, then validation)
split
source split (train / validation)
en_orig
raw BoolQ question (lowercase, no ?, as in google/boolq)
en_preproc
preprocessed input fed to MT: spaCy… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-boolq-questions.EsperantoBenchThis is a benchmark for testing knowledge of the Esperanto language. It contains 651 questions regarding Esperanto vocabulary and grammar. All questions are in 4-answer multiple choice format.
The questions and answers are all based on the book "A Complete Grammar of Esperanto" by Ivy Kellerman Reed. Source: https://www.gutenberg.org/ebooks/7787
As of publishing this on June 23rd, 2025, it seems that this benchmark is already saturated! Neat.
gpt-4.1-nano: 83.87%
gpt-4.1-mini: 93.39%
gpt-4.1:… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/EsperantoBench.
