datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
za-african-next-voices-compressedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate.
Swivuriso: ZA-African Next Voices-Compressed
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.lwazi-asr-corpus-compressed
Lwazi ASR Corpus Collection
This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus for several low-resourced South African languages. These datasets are designed for use in speech recognition research and development, particularly for underrepresented languages.
Corpus Overview
Each corpus consists of scripted telephonic speech recordings collected from native speakers, along with corresponding transcriptions. The… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/lwazi-asr-corpus-compressed.lwazi-asr-corpus-compressed
Lwazi ASR Corpus Collection
This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus for several low-resourced South African languages. These datasets are designed for use in speech recognition research and development, particularly for underrepresented languages.
Corpus Overview
Each corpus consists of scripted telephonic speech recordings collected from native speakers, along with corresponding transcriptions. The… See the full description on the dataset page: https://huggingface.co/datasets/Lindo20/lwazi-asr-corpus-compressed.lwazi-asr-corpus-compressed
Lwazi ASR Corpus Collection
This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus for several low-resourced South African languages. These datasets are designed for use in speech recognition research and development, particularly for underrepresented languages.
Corpus Overview
Each corpus consists of scripted telephonic speech recordings collected from native speakers, along with corresponding transcriptions. The… See the full description on the dataset page: https://huggingface.co/datasets/rareRabbit/lwazi-asr-corpus-compressed.
