kaa
Datasets
All datasets matching “kaa”refusal-exp031-statetsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts
~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz
mono WAV) with systematically repaired transcripts. This is a derivative of
ulaspolat/tsc-tr-filtered-94h,
itself a filtered subset of the ISSAI Turkish Speech Corpus
(MIT license). Audio is unchanged; only the text column was modified.
Transcript repairs
The source transcripts carry two systematic artifacts from İ/apostrophe
mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.KaamelDict
Kaamel-Dict: A Comprehensive Persian G2P Dictionary
Kaamel-Dict is the largest publicly available Persian grapheme-to-phoneme (G2P) dictionary, containing over 116,600 entries.
It was developed by unifying multiple phonetic representation systems from existing G2P tools
[1]
[2]
[3]
[4]
and datasets
[5]
[6]
[7]
[8]
[9]
and an additional online glossary
[10]
, called Jame` Dictionary, . This dictionary is released under the GNU license, allowing it to be used for
developing G2P… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/KaamelDict.UrduTTS
UrduTTS
A Studio-Quality Urdu Speech Corpus with Urdu, Phonemized, and Romanized Transcriptions
91.9 hours · 57,873 utterances · 44.1 kHz · 3 aligned text representations
Dataset Summary
UrduTTS is the largest openly available Urdu text-to-speech corpus with three aligned text
representations for every utterance: native Urdu script, Phonemized (IPA) text, and
Romanized (Latin) text.
Urdu is spoken by roughly 250 million people but is badly… See the full description on the dataset page: https://huggingface.co/datasets/kaab4321/UrduTTS.code2doc
Code2Doc: Function-Documentation Pairs Dataset
A curated dataset of 13,358 high-quality function-documentation pairs extracted from popular open-source repositories on GitHub. Designed for training models to generate documentation from code.
Dataset Description
This dataset contains functions paired with their docstrings/documentation comments from 5 programming languages, extracted from well-maintained, highly-starred GitHub repositories.
Languages Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kaanrkaraman/code2doc.kknews-dataset
KKNews.uz Dataset
Qaraqalpaqstan Xabar Agentligi (kknews.uz) maqalaları — 5 tilde.
Languages
Code
Language
ru
Russian
uz
Uzbek (Latin)
oz
Uzbek (Cyrillic)
kk
Karakalpak (Cyrillic)
qq
Karakalpak (Latin)
Columns
Column
Type
Description
id
int
WordPress post ID
lang
string
Language code
category_id
int
Category ID
category_name
string
Category name
title
string
Plain text title
content_html
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/kknews-dataset.
