datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mzansi-text
MzansiText
MzansiText is a curated multilingual pretraining corpus for all eleven official South African languages.
Dataset details
Splits: 3,943,584 train rows, 19,379 validation rows, and 19,341 test rows
lang values: afr, eng, nbl, nso, sot, ssw, tsn, tso, ven, xho, zul
Schema:
{
"text": "string",
"lang": "string"
}
Validation and test sets are capped at approximately 2M tokens per language to prevent high-resource languages from dominating early… See the full description on the dataset page: https://huggingface.co/datasets/uctnlp/mzansi-text.mzansi-text-deduplicated
MzansiText — deduplicated release
This is a conservatively deduplicated release of MzansiText, a multilingual pretraining corpus for all eleven official South African languages.
Use the original release to reproduce the paper and its trained models. Use this release for new experiments where cross-source duplicate documents should be removed.
Dataset details
Splits: 3,744,654 train rows, 19,940 validation rows, and 19,784 test rows
Total: 3,784,378 rows
lang… See the full description on the dataset page: https://huggingface.co/datasets/uctnlp/mzansi-text-deduplicated.
