CoolFace
Datasetpublic

flexitok/multilingual-addition

Multilingual Addition Dataset Synthetic dataset of addition problems of the form a+b=answer, where a and b are written-form representations of integers in 21 languages, plus a 22nd split using raw digit strings. Task format Each sample contains: field type description a_str str written-form (or digit) representation of a a_digit int integer value of a b_str str written-form (or digit) representation of b b_digit int integer value of b answer str… See the full description on the dataset page: https://huggingface.co/datasets/flexitok/multilingual-addition.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes33downloads
Dataset Card

Multilingual Addition Dataset

Synthetic dataset of addition problems of the form a+b=answer, where a and b are written-form representations of integers in 21 languages, plus a 22nd split using raw digit strings.

Task format

Each sample contains:

fieldtypedescription
a_strstrwritten-form (or digit) representation of a
a_digitintinteger value of a
b_strstrwritten-form (or digit) representation of b
b_digitintinteger value of b
answerstrwritten-form (or digit) of a + b
answer_digitintinteger value of a + b
textstr"{a_str}+{b_str}={answer}" (completion target)
questionstr"{a_str}+{b_str}=" (prompt)
langstrlanguage tag, e.g. eng_Latn, or digit

Numbers range from 0 to 999 for both a and b (answers up to 1998).

Languages

langtrainval
eng_Latn900,000100,000
dan_Latn900,000100,000
swe_Latn900,000100,000
vie_Latn900,000100,000
hun_Latn900,000100,000
fas_Arab900,000100,000
tur_Latn900,000100,000
ces_Latn900,000100,000
arb_Arab900,000100,000
ell_Grek900,000100,000
ind_Latn900,000100,000
nld_Latn900,000100,000
pol_Latn900,000100,000
por_Latn900,000100,000
ita_Latn900,000100,000
jpn_Jpan900,000100,000
fra_Latn900,000100,000
spa_Latn900,000100,000
deu_Latn900,000100,000
cmn_Hani900,000100,000
rus_Cyrl900,000100,000
nob_Latn900,000100,000
fin_Latn900,000100,000
ben_Beng900,000100,000
kor_Hang900,000100,000
digit900,000100,000

Generation

bash
python create_multilingual_addition_data.py \
    --hf_repo_id flexitok/multilingual-addition \
    --publish_to_hf \
    --a_min 0 --a_max 999 \
    --seed 42 --train_ratio 0.9