CoolFace
Datasetpublic

uctnlp/mzansi-text

MzansiText MzansiText is a curated multilingual pretraining corpus for all eleven official South African languages. Dataset details Splits: 3,943,584 train rows, 19,379 validation rows, and 19,341 test rows lang values: afr, eng, nbl, nso, sot, ssw, tsn, tso, ven, xho, zul Schema: { "text": "string", "lang": "string" } Validation and test sets are capped at approximately 2M tokens per language to prevent high-resource languages from dominating early… See the full description on the dataset page: https://huggingface.co/datasets/uctnlp/mzansi-text.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
11likes815downloads
10 commits on main
34026532mo ago

Correct MzansiText raw release (#2)

anrilombard
2ffab246mo ago

Add arXiv paper link and citation

anrilombard
302639c7mo ago

Link dataset cards to GitHub cleaning pipeline

anrilombard
6f6104c7mo ago

Update MzansiText card after adding raw validation and test splits

anrilombard
80ab53b7mo ago

Add raw test split to MzansiText

anrilombard
95651a07mo ago

Add raw validation split to MzansiText

anrilombard
61b381d7mo ago

Refresh MzansiText dataset card

anrilombard
2c7b9d29mo ago

Upload README.md with huggingface_hub

anrilombard
3fb2cb29mo ago

Upload dataset

anrilombard
8499e429mo ago

initial commit

anrilombard