CoolFace
Datasetpublic

simpra/xh-tts-vixsd

isiXhosa TTS clips (ViXSD, segmented) 3,861 clips, 22,050 Hz mono, 8 speakers, cut from long-form recordings by CTC forced alignment. Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI / Way With Words, under the Esethu License — see https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD Pipeline vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous: rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd.

sourceHugging Faceotherupdated 11d agoView on Hugging Face
0likes330downloads
Dataset Card

isiXhosa TTS clips (ViXSD, segmented)

3,861 clips, 22,050 Hz mono, 8 speakers, cut from long-form recordings by CTC forced alignment.

Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI / Way With Words, under the Esethu License — see https://huggingface.co/datasets/lelapa/VukuzenzeleisiXhosaSpeechDatasetViXSD

Pipeline

  1. 1.vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous: rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM and EXTENSIBLE.
  2. 2.vixsd_align.pyfacebook/mms-1b-all + xho adapter, word timestamps.
  3. 3.vixsd_segment.py — cut at sentence > clause > comma > pause, target 6s.

Speakers

speaker_id 0-7. Gender is in metadata.csv; 4 of 8 are male, which is why this dataset exists — SLR32's only male speaker had 7.4 minutes.