CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MaratDV /russian-call-center-speech-ru 📖 Описание на русском ОписаниеКрупный датасет реальных записей колл-центров на русском языке.Телефонное качество, разговоры «клиент–оператор».Подходит для обучения систем ASR (распознавание речи), NLP, голосовых ассистентов и анализа диалогов. Технические характеристики Язык: русский Общая продолжительность: ~832 часа Формат: MP3 Каналы: моно (клиент и оператор в одном канале) Частота дискретизации: 8000 Гц Битрейт: 32 кбит/с Метаданные: отсутствуют… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/russian-call-center-speech-ru.audioautomatic-speech-recognition1 likes114 downloads1y agoHugging Face02ud-nlp /call-center-audio Call Center Dataset - 13,000+ Hours Dataset is a large audio dataset containing 13,000+ hours of real-world customer service calls from global call centers, featuring 90%+ unique speakers and time-stamped transcripts for accurate speech recognition, speaker diarization, and conversational AI model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of real customer service calls Data types Audio Tasks Speech… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/call-center-audio.audioautomatic-speech-recognitionn<1K1 likes54 downloads6mo agoHugging Face03AIxBlock /Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains. 🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions. 🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.audioautomatic-speech-recognitionn<1K5 likes51 downloads1y agoHugging Face04UniDataPro /call-center-audio Call Center Dataset The datasets contain over 13,000+ hours of real-world conversations and feature 90%+ unique speakers. It is designed for advancing conversational AI and speech technologies. The dataset provides high-quality, time-stamped transcripts for model training in speech recognition and speaker diarization, enabling businesses to build and refine their AI systems. By utilizing this dataset, researchers and developers can focus on analyzing customer interactions… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/call-center-audio.audiotext-to-speechn<1K1 likes49 downloads1mo agoHugging Face05InfoBayAI /Arabic-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 64,027 hours of processed Arabic (AR) single-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes43 downloads9d agoHugging Face06InfoBayAI /Hindi_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 1,587,658 hours of processed Hindi dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes42 downloads9d agoHugging Face07InfoBayAI /Nepali_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 229,645 hours of processed Nepalese (NP) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face08InfoBayAI /French_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 31,106 hours of processed French (FR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face09InfoBayAI /Filipino-Tagalog-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4019 hours of processed Tagalog (TL) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino-Tagalog-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face10InfoBayAI /Hindi-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 1,587,658 hours of processed Hindi single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behaviour commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K1 likes40 downloads9d agoHugging Face11InfoBayAI /Arabic_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 64,027 hours of processed Arabic (AR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes39 downloads9d agoHugging Face12InfoBayAI /English-United-Kingdom-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 90,334 hours of processed English (UK) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behavior commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-United-Kingdom-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes39 downloads9d agoHugging Face13InfoBayAI /Russian_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes37 downloads9d agoHugging Face14InfoBayAI /Marathi-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 58,486 hours of processed Marathi (MR) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes36 downloads9d agoHugging Face15InfoBayAI /English_United_States_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 250,362 hours of processed English (US) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_States_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes35 downloads10d agoHugging Face16InfoBayAI /Odia_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 12,794 hours of processed Odia(Oriya) (OR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Odia_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes35 downloads9d agoHugging Face17InfoBayAI /English_United_Kingdom_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 90,334 hours of processed English (UK) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_Kingdom_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K2 likes34 downloads9d agoHugging Face18InfoBayAI /Filipino_Tagalog_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4,019 hours of processed Tagalog (TL) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino_Tagalog_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes34 downloads9d agoHugging Face19InfoBayAI /German_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 212 hours of processed German (DE) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/German_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes32 downloads9d agoHugging Face20InfoBayAI /Mizo_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 468 hours of processed Mizo (MZ) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Mizo_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes31 downloads9d agoHugging Face21InfoBayAI /Russian-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes31 downloads9d agoHugging Face22lilgoose7777 /Nepali_Call_Center_Audio_Dataset_Dual_Channelgated Nepali Call Center Audio Dataset — Dual Channel Source and Attribution This repository contains data obtained from the original dataset published by InfoBay AI Ltd. Original Dataset Name: Nepali_Call_Center_Audio_Dataset_Dual_Channel Original repository: https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel Original creator: InfoBay AI Ltd. The original dataset is listed as CC BY 4.0 on its Hugging Face repository.… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose7777/Nepali_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes31 downloads1mo agoHugging Face23InfoBayAI /English_India_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 30,320 processed English (India) dual-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments.… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_India_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes30 downloads9d agoHugging Face24InfoBayAI /Malayalam-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 14,980 hours of processed Malayalam (ML) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes30 downloads9d agoHugging Face25InfoBayAI /Bengali_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes29 downloads9d agoHugging Face26InfoBayAI /Somali-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 952 hours of processed Somali (SO) and 105 hours of processed Somali (UG) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Somali-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes28 downloads9d agoHugging Face27InfoBayAI /Swahili_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 194,331 hours of processed Swahili (SW) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Swahili_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes27 downloads9d agoHugging Face28InfoBayAI /Spanish_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 452 hours of processed Spanish (MX) and 785 hours of processed Spanish (ES) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Spanish_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes26 downloads9d agoHugging Face29InfoBayAI /Somali_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 952 hours of processed Somali (SO) and 105 hours of processed Somali (UG) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Somali_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes26 downloads10d agoHugging Face30InfoBayAI /Tamil_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 15,056 hours of processed Tamil (TA) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Tamil_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes26 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.