McGill-NLP/NaijaS2ST
Multilingual Speech Dataset (Speech-to-Speech / Speech-to-Text Ready) *The IWSLT shared task submission details and the test set are now available at IWSLT 2026 * Dataset Summary This dataset is a large-scale multilingual speech corpus curated for speech-to-speech translation, speech-to-text, and multilingual speech processing research. The data is organized by language, speaker (user_id), and dataset split (train, dev), and includes rich acoustic and metadata… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/NaijaS2ST.
This repository belongs to McGill-NLP on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
