CoolFace
Datasetpublic

Adeptschneider/CiviVox-Swahili-text-corpus-v2.0

Swahili Text Dataset Overview This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Dataset Details Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset Content The dataset consists of two main columns: id: A unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/Adeptschneider/CiviVox-Swahili-text-corpus-v2.0.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes7downloads
settings

This repository belongs to Adeptschneider on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCiviVox-Swahili-text-corpus-v2.0
visibilitypublic
licenceapache-2.0
gatedno
ownerAdeptschneider
Account settings
Adeptschneider/CiviVox-Swahili-text-corpus-v2.0 · CoolFace