datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
za-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.multilingual-nchlt-dataset
NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset
Dataset Description
This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research.
The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.anv-kikuyu-banking-subset
Anv-Kikuyu Banking Subset
A domain-filtered subset of Kikuyu (Gĩkũyũ) speech data focused on banking and financial-transaction content, combined into a single repository with train, test, and validation splits.
Source
This dataset is a filtered subset of Anv-ke/kikuyu, part of the African Next Voices (ANV) collection. All audio, transcriptions, and underlying speaker data originate from that source dataset. Full credit for data collection belongs to the… See the full description on the dataset page: https://huggingface.co/datasets/NjeriKahoro/anv-kikuyu-banking-subset.za-african-next-voices-compressedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate.
Swivuriso: ZA-African Next Voices-Compressed
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.anv-za-sot-1h-sample-dataset
Sesotho Sample Dataset - Next Voices-ZA (South Africa) - Multilingual Speech Dataset - Sesotho
This dataset includes scripted and unscripted speech across various domains such as agriculture, health, finance, sports, transport, culture, society and general topics. It is primarily designed for automatic speech recognition (ASR).
Use Restriction:
The persons whose voices are included in this dataset, and the creators and owners of this dataset* do not give consent in… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/anv-za-sot-1h-sample-dataset.
