CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Taklaxbr /turkish-tts-combined-raw Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw 🔗 Derleyen Platform: VeriPazarı Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined) 7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir. ~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K0 likes792 downloads3mo agoHugging Face02SaarAI /waxal-amharic-combinedaudio100K<n<1M0 likes553 downloads1mo agoHugging Face03projectkaira /turkish-tts-combined-raw Türkçe TTS Birleşik Veri Seti 7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu Kaynaklar Veri Seti Örnek Kaynak Mazlum Kiper 9,643 omersaidd/tts_mazlum_kiper_tur Ahmet Deniz 11,289 omersaidd/tts_ahmet_deniz_tur Nisan Kumru 8,042 omersaidd/tts_nisan_kumru_tur Derya TTS v2 42 afkfatih/derya-tts-v2 Derya Karma v3 255 afkfatih/derya-tts-karma-v3 Khan Academy 25,741 ysdede/khanacademy-turkish Common… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K1 likes494 downloads2mo agoHugging Face04afkfatih /turkish-tts-combined-raw Türkçe TTS Birleşik Veri Seti 7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu Kaynaklar Veri Seti Örnek Kaynak Mazlum Kiper 9,643 omersaidd/tts_mazlum_kiper_tur Ahmet Deniz 11,289 omersaidd/tts_ahmet_deniz_tur Nisan Kumru 8,042 omersaidd/tts_nisan_kumru_tur Derya TTS v2 42 afkfatih/derya-tts-v2 Derya Karma v3 255 afkfatih/derya-tts-karma-v3 Khan Academy 25,741 ysdede/khanacademy-turkish Common Voice 17 26… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K15 likes392 downloads10mo agoHugging Face05AbstractTTS /combined_dataset_mapaudio100K<n<1M2 likes338 downloads2y agoHugging Face06Razer112 /Expresso-combined Disclaimer This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility. audion<1K0 likes325 downloads2mo agoHugging Face07Razer112 /Hifi-TTS-combined Disclaimer This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility. audion<1K0 likes325 downloads2mo agoHugging Face08barro /waxal-asr-lin_sna_lug-combined-datasetaudio10K<n<100K0 likes296 downloads2mo agoHugging Face09rootflo /combined_v4audio100K<n<1M0 likes269 downloads2y agoHugging Face10windcrossroad /combined_dataset_mapaudio100K<n<1M0 likes263 downloads2y agoHugging Face11Razer112 /m4singer-combined Disclaimer This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility. audion<1K0 likes228 downloads2mo agoHugging Face12asr-malayalam /combined_malayalamaudio10K<n<100K0 likes213 downloads2y agoHugging Face13nickfuryavg /combined-audio-datasetaudio1K<n<10K0 likes208 downloads1y agoHugging Face14chosenek /czech-speech-combined Czech Speech Combined Dataset Quality-filtered Czech speech dataset for TTS/ASR training. 146,153 clips across ~1,700 speakers from 5 sources. Sources Source Clips Hours Speakers Origin audiobooks 24,535 ~33h 13 Czech audiobook narrations audiobooks_new 28,131 ~39h 8+ Czech audiobook narrations yodas_czech 28,208 ~37h ~2,300 YODAS YouTube speech (quality-filtered) voxpopuli_czech 12,679 ~33h 45 VoxPopuli parliament speech commonvoice_czech 52… See the full description on the dataset page: https://huggingface.co/datasets/chosenek/czech-speech-combined.text100K<n<1M2 likes170 downloads4mo agoHugging Face15Shiry /ATC_combined Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Shiry/ATC_combined.audioautomatic-speech-recognition10K<n<100K1 likes165 downloads3y agoHugging Face16Hm12cbbcbx /turkish-tts-combined-raw Türkçe TTS Birleşik Veri Seti 7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu Kaynaklar Veri Seti Örnek Kaynak Mazlum Kiper 9,643 omersaidd/tts_mazlum_kiper_tur Ahmet Deniz 11,289 omersaidd/tts_ahmet_deniz_tur Nisan Kumru 8,042 omersaidd/tts_nisan_kumru_tur Derya TTS v2 42 afkfatih/derya-tts-v2 Derya Karma v3 255 afkfatih/derya-tts-karma-v3 Khan Academy 25,741 ysdede/khanacademy-turkish Common… See the full description on the dataset page: https://huggingface.co/datasets/Hm12cbbcbx/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K2 likes162 downloads3mo agoHugging Face17Gbssreejith /combined_malyalam_voice_datasetaudio10K<n<100K0 likes145 downloads2y agoHugging Face18Eimhin03 /combined_3datasets_801010audio1K<n<10K0 likes140 downloads6mo agoHugging Face19Tabys /ATC_combined Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Tabys/ATC_combined.audioautomatic-speech-recognition10K<n<100K0 likes128 downloads7mo agoHugging Face20timniel /Pidgin_ASR_Dataset_Combined Naija-ASR-Corpus v1.0 (NAC-v1.0) A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM) 📌 Dataset Summary Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC). The NAC Team processed the original long-form recordings by: Segmenting the audio into sentence-level clips. Transcribing/Aligning the text to create paired audio-text data suitable for ASR training. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/timniel/Pidgin_ASR_Dataset_Combined.audioautomatic-speech-recognition1K<n<10K0 likes124 downloads10mo agoHugging Face21KYAGABA /COMBINED_TEST_AMHARICaudio1K<n<10K0 likes122 downloads2y agoHugging Face22averageandyyy /combined_librispeech_self Dataset Card for "combined_librispeech_self" num_examples: 2620 (test) num_examples: 28539 (train) audio10K<n<100K0 likes121 downloads3y agoHugging Face23AnonXx /Pidgin_ASR_Dataset_Combined Naija-ASR-Corpus v1.0 (NAC-v1.0) A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM) 📌 Dataset Summary Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC). The NAC Team processed the original long-form recordings by: Segmenting the audio into sentence-level clips. Transcribing/Aligning the text to create paired audio-text data suitable for ASR training. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/AnonXx/Pidgin_ASR_Dataset_Combined.audioautomatic-speech-recognition1K<n<10K1 likes114 downloads6mo agoHugging Face24Codyfederer /tr-combinedgated Tr_combined This is a merged speech dataset containing 221531 audio segments from 894 source datasets. Dataset Information Total Segments: 221531 Speakers: 2158 Languages: tr Emotions: happy, angry, neutral, sad Original Datasets: 894 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) language:… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-combined.audioautomatic-speech-recognition100K<n<1M6 likes109 downloads1y agoHugging Face25IbrahimDayax /somali-combined-asr-stt-dataset Somali Combined ASR/STT Dataset A Somali automatic speech recognition (ASR) / speech-to-text (STT) dataset combining synthetic TTS-generated audio and other Somali speech sources, deduplicated by transcript text and split into train/validation/test. Dataset Summary Language: Somali (so) Task: Automatic Speech Recognition / Speech-to-Text Audio format: WAV, 16 kHz mono Total examples: 8,226 (after deduplication) Total audio: ~6 hours Split Examples Audio… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-combined-asr-stt-dataset.audioautomatic-speech-recognition1K<n<10K0 likes107 downloads2mo agoHugging Face26KYAGABA /combined_amharic_speech_datasetaudio100K<n<1M0 likes105 downloads2y agoHugging Face27michaelodafe /pidgin-asr-combined Pidgin ASR Combined A unified Nigerian Pidgin English speech-to-text dataset that combines publicly available Pidgin ASR sources into a single train / validation / test setup with a consistent schema. Built for fine-tuning Whisper-family models on Nigerian Pidgin (Naija, pcm). ~8.6 hours, 4,278 clips, 10 source speakers, 16 kHz mono WAV. Used to train michaelodafe/whisper-pidgin-v1 (21.37% WER on the test split, beating the published Wav2Vec2-XLSR-53 baseline by 8.2 pp).… See the full description on the dataset page: https://huggingface.co/datasets/michaelodafe/pidgin-asr-combined.audioautomatic-speech-recognition1K<n<10K1 likes104 downloads5mo agoHugging Face28youssefkhalil320 /MedSynth-Combinedaudio1K<n<10K0 likes76 downloads5mo agoHugging Face29Eimhin03 /combined-irish-asr-801010audio1K<n<10K0 likes75 downloads7mo agoHugging Face30Eimhin03 /combined-irish-asr-801010-no-fleursaudio1K<n<10K0 likes66 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.