datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indommlu-local-languages
IndoMMLU: Local Languages and Cultures (audited subset)
An audited, corrected subset of IndoMMLU
(Koto et al., 2023) covering the 9 Local Languages and Cultures subjects:
Indonesian primary and secondary school exam questions written in Balinese,
Banjarese, Dayak Ngaju, Javanese, Lampung, Madurese, Makassarese, and
Sundanese, plus one culture-knowledge subject on Minangkabau customs
(answered in standard Indonesian). This is not a dataset we created.
It is IndoMMLU's own subset… See the full description on the dataset page: https://huggingface.co/datasets/ibahasa/indommlu-local-languages.gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.proxy-mt-benchmark-scores
Proxy-MT Benchmark Scores
Multilingual benchmark results for 50 open-weight LLMs, evaluated with the
lm-evaluation-harness via a vLLM
backend. Covers reasoning, comprehension, and knowledge tasks with an emphasis on
African and other lower-resource languages.
Layout
scores/<model>.csv # parsed per-language scores (tidy, ready to plot)
raw/<model>/.../results_*.json # raw lm-eval-harness result files
raw/<model>/raw_log.txt # full evaluation… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-benchmark-scores.
