CoolFace
Datasetpublic

NCSpeech/YO-CPT-kk

YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
10likes1.1kdownloads
settings

This repository belongs to NCSpeech on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameYO-CPT-kk
visibilitypublic
licenceother
gatedno
ownerNCSpeech
Account settings
NCSpeech/YO-CPT-kk · CoolFace