simpra/xh-tts-vixsd
isiXhosa TTS clips (ViXSD, segmented) 3,861 clips, 22,050 Hz mono, 8 speakers, cut from long-form recordings by CTC forced alignment. Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI / Way With Words, under the Esethu License — see https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD Pipeline vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous: rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd.
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI / Way With Words, under the Esethu License — see https://huggingface.co/datasets/lelapa/VukuzenzeleisiXhosaSpeechDatasetViXSD
Pipeline
vixsd_extract.py— parquet to mono 22,050 Hz. Source is heterogeneous: rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM and EXTENSIBLE.vixsd_align.py—facebook/mms-1b-all+xhoadapter, word timestamps.vixsd_segment.py— cut at sentence > clause > comma > pause, target 6s.
Speakers
speaker_id 0-7. Gender is in metadata.csv; 4 of 8 are male, which is why this dataset exists — SLR32's only male speaker had 7.4 minutes.
