datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xh-tts-vixsd
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd.xh-tts-vixsd-norm
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd-norm.Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Dataset Card for Vuk'uzenzele isiXhosa Speech Dataset (ViXSD)
Dataset Description
Dataset Summary
Vuk'uzenzele isiXhosa Speech Dataset (ViXSD) contains scripted narration of the Vuk’uzenzele South African Multilingual Corpus.
ViXSD contains read speech from native speakers accompanied with rich metadata on speaker demographic and linguistic distribution.
ViXSD consists of 395 stereo audio recordings and corresponding transcriptions derived from the… See the full description on the dataset page: https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD.
