CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01InfoBayAI /Tamil_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 15,056 hours of processed Tamil (TA) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Tamil_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes29 downloads11d agoHugging Face02InfoBayAI /Tamil-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 15,056 hours of processed Tamil (TA) Single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Tamil-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes25 downloads10d agoHugging Face03InfoBayAI /Tamil_Podcast_Audio_DatasetgatedDataset Description: This dataset is a large-scale collection of 3,315 hours of processed Tamil podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It captures real-world interactions across diverse topics and formats. The dataset preserves natural speech patterns, speaker variability, and authentic podcast environments, making it highly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Tamil_Podcast_Audio_Dataset.audioautomatic-speech-recognitionn<1K1 likes21 downloads10d agoHugging Face04InfoBayAI /Tamil_Podcast_Audio_Dataset_Dual_Channelgated Dataset Description This dataset is a large-scale collection of 3,315 hours of processed Tamil dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker audio… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Tamil_Podcast_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes18 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.