KaanAydinli/tsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts ~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz mono WAV) with systematically repaired transcripts. This is a derivative of ulaspolat/tsc-tr-filtered-94h, itself a filtered subset of the ISSAI Turkish Speech Corpus (MIT license). Audio is unchanged; only the text column was modified. Transcript repairs The source transcripts carry two systematic artifacts from İ/apostrophe mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.
TSC-TR Filtered 94h — repaired transcripts
~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz mono WAV) with systematically repaired transcripts. This is a derivative of ulaspolat/tsc-tr-filtered-94h, itself a filtered subset of the ISSAI Turkish Speech Corpus (MIT license). Audio is unchanged; only the text column was modified.
Transcript repairs
The source transcripts carry two systematic artifacts from İ/apostrophe mishandling upstream: split dotted-capital İ words (i stanbul, i ngilizce) and orphaned apostrophe suffixes (kavala nın, adana dan, haber e). 8,402 lines (11.6%) were repaired with 9,676 merges, guarded by corpus token frequency (a fragment like stanbul never occurs as a free token) and Turkish vowel harmony (2-way for -dan/-a, 4-way for -nın/-yı; vowel-less acronyms like lgbt yi pass freely):
Ambiguous free tokens (de/da/ya/la — real conjunctions) were never merged. 412 lone vowels with no safe merge (mostly foreign names whose written vowels break harmony, e.g. trump ı) were left untouched. Transcripts remain lowercase and unpunctuated as in the source.
Columns
audio— WAV, 16 kHz monotext— repaired transcript (lowercase, no punctuation)source— pseudo speaker label from the source dataset's clustering
Credits and citation
Full credit for the audio and original transcripts belongs to ISSAI (Nazarbayev University); the filtering (noise removal, speaker clustering) is by ulaspolat. Please cite the original corpus:
@Article{info14020074,
AUTHOR = {Mussakhojayeva, Saida and Dauletbek, Kaisar and Yeshpanov, Rustem and Varol, Huseyin Atakan},
TITLE = {Multilingual Speech Recognition for Turkic Languages},
JOURNAL = {Information},
VOLUME = {14},
YEAR = {2023},
NUMBER = {2},
ARTICLE-NUMBER = {74},
URL = {https://www.mdpi.com/2078-2489/14/2/74},
ISSN = {2078-2489},
DOI = {10.3390/info14020074}
}