KaanAydinli/tsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts ~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz mono WAV) with systematically repaired transcripts. This is a derivative of ulaspolat/tsc-tr-filtered-94h, itself a filtered subset of the ISSAI Turkish Speech Corpus (MIT license). Audio is unchanged; only the text column was modified. Transcript repairs The source transcripts carry two systematic artifacts from İ/apostrophe mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.
0577
Upload README.md with huggingface_hub
Upload dataset
initial commit
