datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hypa-Voices-snac
Hypa-Voices-snac
This repository is the SNAC token companion to hypaai/Hypa-Voices.
Both datasets belong to the Hypa-Voices collection and contain the same 8,800 curated records with identical metadata columns. The only difference is how speech is stored:
Repository
Speech field
Format
hypaai/Hypa-Voices
audio
FLAC (decoded waveform)
hypaai/Hypa-Voices-snac (this repo)
codes_list
SNAC discrete token sequence
For the full dataset description, data fields… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Voices-snac.kinyarwanda-pastor-snac
Kinyarwanda Pastor Uwambaje — SNAC Dataset
Single-speaker Kinyarwanda TTS dataset from Pastor Uwambaje YouTube sermons.
Audio cleaned with htdemucs (music removal) + resemble-enhance (denoising).
Encoded with SNAC 24kHz, 7-token interleaved format.
Total uploaded: 13,380 clips at STOI >= 0.80 threshold (~18.4h).
STOI Quality Distribution
STOI Threshold
Clips
Hours
>= 0.80 (this dataset)
13,225
18.4h
>= 0.85
13,058
18.2h
>= 0.90
12,488
17.5h
>= 0.95
9,734… See the full description on the dataset page: https://huggingface.co/datasets/vysakh25/kinyarwanda-pastor-snac.
