datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts
~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz
mono WAV) with systematically repaired transcripts. This is a derivative of
ulaspolat/tsc-tr-filtered-94h,
itself a filtered subset of the ISSAI Turkish Speech Corpus
(MIT license). Audio is unchanged; only the text column was modified.
Transcript repairs
The source transcripts carry two systematic artifacts from İ/apostrophe
mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.jo-asrTraditionalDataset4v5Qualitykarakalpak-audio-datasetNewDatasetv2IEEEAccessDatasetSLRVoiceskllm_dialogTraditionalDataset4v3filtered_voicesbenchmark-rawsophiamerged_urdu_TTScommon_voice_filteredTraditionalDataset4v4IEEEAccessDatasetTraditionalDataset4indic_voices_filteredtts_filteredTraditionalDataset4v1ow-asr-test-dataSlrCvVoicesTtsDatasetNewDatasetv1fleursCommonVoice17-Clonebenchmark-querieskllm_storykllm-bench-queriesflrs_with_top10_contentkllm_top10
