datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnilora-kazakh-child-mvp
OmniLoRA Kazakh Child-Voice TTS — Cleaned & Emotion-Labeled Subset
A 535-clip Kazakh child-speech subset derived from
galammadin-asr/child-asr-kazakh,
cleaned through a 3-stage automatic filter and hand-labeled with one of six
emotion categories. Built for fine-tuning a LoRA adapter on top of
OmniVoice (Method 5 of a 5-method Kazakh
TTS benchmark, CSCI 595 final project).
Dataset summary
Clips
535
Language
Kazakh (kk)
Sample rate
16 kHz (inherits from… See the full description on the dataset page: https://huggingface.co/datasets/neversi123/omnilora-kazakh-child-mvp.sep28k-mvpHeimdall_MVP_DataSet
