datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kazlibriKazLibri
KazLibri is a Kazakh narrated speech corpus consisting of audio–text pairs derived from books. The project is inspired by the LibriSpeech corpus.
The corpus is released strictly for research and educational purposes only. Any form of commercial use, redistribution for profit, or integration into proprietary systems is not permitted. At the time of public release, the corpus contains just over six hours of narrated speech.
Initiated independently during a period of unemployment… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/kazlibri.kazakh_songs_asr
Kazakh Songs ASR Dataset
Dataset Summary
This dataset consists of manually aligned audio–text pairs extracted from Kazakh songs and designed for research in automatic speech recognition (ASR) for low-resource languages. The primary goal of the dataset is to investigate whether sung speech can serve as a complementary training resource for Kazakh ASR systems.
The corpus contains line-level vocal segments obtained from commercially released songs, with manually verified… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/kazakh_songs_asr.
