CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SPRINGLab /IndicTTS_Telugu Telugu Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Telugu monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Telugu Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Telugu.audiotext-to-speech1K<n<10K7 likes423 downloads2y agoHugging Face02Noothi /telugu-tech-custom-voice 🎙️ Telugu Tech Custom Voice Dataset A high-quality, clean single-speaker Telugu Speech & Voice dataset tailored for training and fine-tuning neural Text-to-Speech (TTS) models (e.g. Coqui XTTS v2, Piper TTS, VITS, Bark) and Automatic Speech Recognition (ASR). 📊 Dataset Statistics Total Clips: 455 audio files (.wav) Total Audio Duration: 1 Hour 12 Minutes 48.5 Seconds (4,368.5 seconds) Total Dataset Size: ~1.20 GB Language: Telugu (te) with technical terms /… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-custom-voice.audiotext-to-speechn<1K0 likes248 downloads2mo agoHugging Face03eswardivi /Telugu_ASR_corpus Dataset Card for "Telugu_ASR_corpus" More Information needed audio1K<n<10K1 likes246 downloads3y agoHugging Face04zsy12345 /telugu-asraudio100K<n<1M2 likes119 downloads3y agoHugging Face05Noothi /telugu-tech-indicf5-custom-voice 🎙️ Telugu Tech IndicF5 Custom Voice Dataset A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models. All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.audiotext-to-speech100K<n<1M0 likes80 downloads2mo agoHugging Face06arpit-tiwari /syspin-telugu-ttsaudio10K<n<100K1 likes79 downloads1y agoHugging Face07shunyalabs /telugu-speech-datasetaudio10K<n<100K3 likes78 downloads1y agoHugging Face08saurabh1043 /noisy_teluguaudio100K<n<1M0 likes77 downloads4mo agoHugging Face09Noothi /telugu-indicf5-evaluationaudion<1K0 likes66 downloads2mo agoHugging Face10sreesaiarjun /telugu-voices-rawaudio0 likes56 downloads3mo agoHugging Face11deboleen6 /telugu_OpenSLRaudio1K<n<10K0 likes53 downloads2y agoHugging Face12Sonal0205 /telugu_whisper_asraudio1K<n<10K0 likes51 downloads2y agoHugging Face13Sonal0205 /Telugu_whisper_ASR_datasetaudioautomatic-speech-recognition1K<n<10K1 likes35 downloads2y agoHugging Face14KarthikeyaPranav /telugu_asr_630hraudio100K<n<1M1 likes35 downloads16d agoHugging Face15kattojuprashanth238 /Telugu-Audio-Corpusaudion<1K1 likes33 downloads2y agoHugging Face1634data /indic-tts-teluguaudio1K<n<10K0 likes32 downloads2mo agoHugging Face17chsasank /tts_synthetic_te-IN_Teluguaudio1K<n<10K1 likes29 downloads1y agoHugging Face18VistaCruzer /telugu-movies-speechaudion<1K0 likes22 downloads2y agoHugging Face19saurabh1043 /telugu-raw-audioaudio10K<n<100K0 likes22 downloads9mo agoHugging Face20InfoBayAI /Telugu_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 27,787 hours of processed Telugu (TE) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Telugu_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes22 downloads10d agoHugging Face21tmtanu /TTS_Telugu Telugu Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Telugu monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Telugu Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours) Audio Format: WAV Sampling… See the full description on the dataset page: https://huggingface.co/datasets/tmtanu/TTS_Telugu.audiotext-to-speech1K<n<10K0 likes21 downloads4mo agoHugging Face22InfoBayAI /Telugu-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 27,787 hours of processed Telugu (TE) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Telugu-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes20 downloads10d agoHugging Face23jigsaws-stomper /telugu_v2_evalaudion<1K0 likes20 downloads3mo agoHugging Face24JaiJaya /telugu-asr-speech-dataaudio10K<n<100K1 likes19 downloads1y agoHugging Face25jigsaws-stomper /telugu_v3_evalaudion<1K0 likes19 downloads3mo agoHugging Face26dianavdavidson /Vaani-telugu-lg-English-no-transcript1audio1K<n<10K0 likes17 downloads4mo agoHugging Face27InfoBayAI /Telugu_Podcast_Audio_Dataset_Dual_Channelgated Dataset Description This dataset is a large-scale collection of 5,964 hours of processed Telugu dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI, automatic speech recognition (ASR), speaker understanding, audio analytics, and multilingual language technologies. It captures real-world podcast conversations across diverse topics and… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Telugu_Podcast_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes17 downloads10d agoHugging Face28Manivarma /ai4bharat_telugu_datasetaudio100K<n<1M2 likes16 downloads1y agoHugging Face29dianavdavidson /Vaani-telugu-lg-English-no-transcript0audio10K<n<100K0 likes16 downloads4mo agoHugging Face30Akshayakrishna /telugu-speech-datasetaudio10K<n<100K0 likes16 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.