naija
Datasets
All datasets matching “naija”NaijaS2ST
Multilingual Speech Dataset (Speech-to-Speech / Speech-to-Text Ready)
*The IWSLT shared task submission details and the test set are now available at IWSLT 2026 *
Dataset Summary
This dataset is a large-scale multilingual speech corpus curated for speech-to-speech translation, speech-to-text, and multilingual speech processing research.
The data is organized by language, speaker (user_id), and dataset split (train, dev), and includes rich acoustic and metadata… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/NaijaS2ST.naija3naija2naijavoices-dataset
Important Information: please be aware that this version of the dataset is huge (500+ GB) and can therefore be challenging to use. To alleviate this and facilitate adoption, we’ve provided a compressed version (84GB) here. We suggest using that instead if you have compute/storage constraints. They are both the exact same data.
Introduction
Welcome to the NaijaVoices dataset. The NaijaVoices dataset consists of 1,800 hours of authentic speech (from over 5,000 diverse speakers!)… See the full description on the dataset page: https://huggingface.co/datasets/naijavoices/naijavoices-dataset.African_voices_naija
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Nigerian Pidgin (pcm) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Nigerian Pidgin (pcm). This corpus is designed to accelerate the development of speech technology in African contexts, promoting… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_naija.naija-speech-afrispeech-ng
