CoolFace
Datasetpublic

uctnlp/mzansi-text-deduplicated

MzansiText — deduplicated release This is a conservatively deduplicated release of MzansiText, a multilingual pretraining corpus for all eleven official South African languages. Use the original release to reproduce the paper and its trained models. Use this release for new experiments where cross-source duplicate documents should be removed. Dataset details Splits: 3,744,654 train rows, 19,940 validation rows, and 19,784 test rows Total: 3,784,378 rows lang… See the full description on the dataset page: https://huggingface.co/datasets/uctnlp/mzansi-text-deduplicated.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes629downloads

uctnlp/mzansi-text-deduplicated · main · files are served by the source, never re-hosted here