CoolFace
Datasetpublic

rlabz/quantum-tts-tokenized

Swahili (swa_spk3) SNAC-Tokenized Dataset for Orpheus-TTS Fine-Tuning Dataset Summary A single-speaker Kiswahili subset, resampled and tokenized for fine-tuning Orpheus-TTS. It is derived from rlabz/swa_lug_tts by: Filtering the train and validation splits down to speaker swa_spk3 only. Resampling all audio from its original 22,050 Hz to 24,000 Hz, the sample rate required by SNAC (snac_24khz), the neural audio codec Orpheus is trained on. Encoding each clip with… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/quantum-tts-tokenized.

sourceHugging Facecc0-1.0updated 1mo agoView on Hugging Face
0likes207downloads
../
filetrain-00000-of-00001.parquet2.8 MBdownload

rlabz/quantum-tts-tokenized · main · files are served by the source, never re-hosted here