datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CiviVox-Swahili-text-corpus-v2.0
Swahili Text Dataset
Overview
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Dataset Details
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: 1.54M
Format: Hugging Face Dataset
Content
The dataset consists of two main columns:
id: A unique identifier for each text entry
text:… See the full description on the dataset page: https://huggingface.co/datasets/Adeptschneider/CiviVox-Swahili-text-corpus-v2.0.CiviVox-Swahili-text-corpus-v2.0
Swahili Text Dataset
Overview
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Dataset Details
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: 1.54M
Format: Hugging Face Dataset
Content
The dataset consists of two main columns:
id: A unique identifier for each text entry
text:… See the full description on the dataset page: https://huggingface.co/datasets/niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0.
