datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.YO-CPT-kk
YO-CPT-kk
YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily
quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker,
TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a
punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and
cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the
voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.kmr-bibleQuechua_datasetdatapruebaQuechua_Spanish_datasetdataset_cppushto-speech-to-text-dataset___CS224S_Quechua_Projectsortformer-synth-coral-corpus-a6a28fcpmarselocp-final-project2CS224S_Quechua_Project_emotion_datasetpp-attach-audiocp-final-projectCS224S_Spanish_subsetfacebook_mms-tts-eng_GPU-CPUjenny_final_selection_audioinference_long_DAC-SE2_np_v5_cp7k_TESTparliament_14263qwen3-tts-tamil-cpt-samples
Qwen3-TTS Tamil CPT - Audio Samples
Step-50 Checkpoint vs Base Model
Generated with voice cloning from a Tamil reference audio.
Files
checkpoint_ta_*.wav - Step-50 checkpoint, Tamil prompts
checkpoint_en_*.wav - Step-50 checkpoint, English prompts
base_ta_*.wav - Base model (untrained), Tamil prompts
base_en_*.wav - Base model (untrained), English prompts
ref_audio.wav - Reference audio used for voice cloning
Results Summary
Model
Tamil EOS… See the full description on the dataset page: https://huggingface.co/datasets/gakashguru/qwen3-tts-tamil-cpt-samples.drzdst
