datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mgsm_gl
Dataset Card for mgsm_gl
mgsm_gl is a question answering dataset in Galician translated from the MGSM dataset in English.
Dataset Details
Dataset Description
This dataset is the Galician version of the MGSM (Multilingual Grade School Math) dataset. It serves as a benchmark of grade-school math problems as proposed in the paper Language models are multilingual chain-of-thought reasoners. It includes 8 instances in the train split and another 250 instances in the… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/mgsm_gl.llm-metric-mgsm2026-01-01-mgsm-subset-en-es-fr-de
2026-01-01-mgsm-subset-en-es-fr-de
Input dataset for the logprob_pivots track: a fixed 50-problems-per-language MGSM subset so all runs score the same problems.
field
value
date_generated
2026-01-01
track
logprob_pivots
languages
en,es,fr,de
models
none (input data; derived from juletxara/mgsm)
provenance
python -m scripts.logprob_pivots.build_mgsm_subset (seed 42, n_per_lang 50, src/logprob_pivots/config.py)
schema
CSV: one row per problem x language with… See the full description on the dataset page: https://huggingface.co/datasets/multicot/2026-01-01-mgsm-subset-en-es-fr-de.mgsmflan-t5-boosting-mgsm_cotflan-t5-boosting-mgsm_zsmgsmupdatetest
mgsm_equationsmgsm_questions
