datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
za-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.za-african-next-voices-compressedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate.
Swivuriso: ZA-African Next Voices-Compressed
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.za-african-next-voices-tonal
Tonal Dataset for Bantu Languages
This dataset contains tonal (F0/pitch) metadata extracted from dsfsi-anv/za-african-next-voices.
Languages
zul
xho
sot
tsn
ven
tso
Files Structure
tonal_data/
<split>/ # train, dev_test, etc.
<lang>/
utterance_tonal_stats.csv # Per-utterance tonal statistics
f0_syllables.csv # Raw F0 segments (word-level)
f0_syllables_clustered.csv # F0 segments with tone cluster labels… See the full description on the dataset page: https://huggingface.co/datasets/kesbeast23/za-african-next-voices-tonal.za-african-next-voices-difficulty-scoresza-african-next-voices-difficulty-scores-5za-african-next-voices-tonal-local
