kazakh
Datasets
All datasets matching “kazakh”Kazakh_Speech_Corpus_2
Kazakh Speech Corpus 2 (KSC2)
This dataset card describes the KSC2, an industrial-scale, open-source speech corpus for the Kazakh language.
Paper: KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus
Summary: KSC2 corpus subsumes the previously introduced two corpora: Kazakh Speech Corpus and Kazakh Text-To-Speech 2, and supplements additional data from other sources like tv programs, radio, senate, and podcasts. In total, KSC2 contains around 1.2k hours of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2.Kazakh-Russian-Child-Directed-Speech-Corpus
Kazakh-Russian Child-Directed Language Corpus
Dataset Description
The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization.
The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts.
This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.multidomain-kazakh-dataset
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/multidomain-kazakh-dataset.ISSAI_KazakhTTS2kazakh-ocr
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
Synthetic OCR benchmark for Arabic, Cyrillic, and Latin script Kazakh OCR. Labels are in metadata.csv.
Curated by: Henry Gagnier, Sophie Gagnier, Ashwin Kirubakaran
Language(s) (NLP): Kazakh (kk)
License: MIT
Citation
@inproceedings{gagnier-etal-2026-kazakhocr,
title = "{K}azakh{OCR}: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource {K}azakh… See the full description on the dataset page: https://huggingface.co/datasets/henrygagnier/kazakh-ocr.kazakh-traditional-audio
Kazakh Traditional Audio Archive
🇰🇿 Қазақша
Бұл жинақ – қазақтың дәстүрлі дыбыстық мұрасын сақтау және машиналық оқыту модельдерін дамытуға арналған ашық дерекқор.
Мазмұны:
Күйлер (инструменталды)
Жыр-термелер (поэзия, әнмен оқу)
Халық әндері және халық композиторларының шығармалары
Ертегілер (аудио нұсқада)
Бесік жырлары
Ретро әндер (1950–1990)
Форматтар:
WAV (PCM), 44.1 кГц
MOНО
metadata.jsonl арқылы сипаттама берілген
Қолдану… See the full description on the dataset page: https://huggingface.co/datasets/rtrk/kazakh-traditional-audio.
