CoolFace
Datasetpublic

KaanAydinli/tsc-tr-filtered-94h-clean

TSC-TR Filtered 94h — repaired transcripts ~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz mono WAV) with systematically repaired transcripts. This is a derivative of ulaspolat/tsc-tr-filtered-94h, itself a filtered subset of the ISSAI Turkish Speech Corpus (MIT license). Audio is unchanged; only the text column was modified. Transcript repairs The source transcripts carry two systematic artifacts from İ/apostrophe mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes577downloads
3 commits on main
96372f01mo ago

Upload README.md with huggingface_hub

KaanAydinli
3c22b901mo ago

Upload dataset

KaanAydinli
10db3be1mo ago

initial commit

KaanAydinli