datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afriscience_mt
AfriScience-MT
A parallel scientific machine-translation corpus for English + six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, isiZulu), co-developed with expert science communicators and professional translators across 11 scientific domains (Agriculture, Biochemistry, Biology, Chemistry, Computer Science, Engineering, Geography, Health, Indigenous Knowledge, Sociology, Statistics).
Alongside the corpus we release every model prediction and per-run metric… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/afriscience_mt.afriscience_mt
AfriScience-MT
A parallel scientific machine-translation corpus for English + six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, isiZulu), co-developed with expert science communicators and professional translators across 11 scientific domains (Agriculture, Biochemistry, Biology, Chemistry, Computer Science, Engineering, Geography, Health, Indigenous Knowledge, Sociology, Statistics).
Alongside the corpus we release every model prediction and per-run metric… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afriscience_mt.
