CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01apptek-com /apptek_callcenter_dialogues AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions. 128.6 hours of speech 14 English accent groups 16 service domains 5–15 minute conversations (long-form) Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.audioautomatic-speech-recognition1K<n<10K37 likes2.4k downloads1mo agoHugging Face02markmuller /call-center-prod-datagated1 likes2.1k downloads4h agoHugging Face03AIxBlock /92k-real-world-call-center-scripts-englishArXiv Paper Publication Here: "Real-World En Call Center Transcripts Dataset with PII Redaction" This dataset includes 91,706 high-quality transcriptions corresponding to approximately 10,500 hours of real-world call center conversations in English, collected across various industries and global regions. The dataset features both inbound and outbound calls and spans multiple accents, including Indian, American, and Filipino English. All transcripts have been carefully redacted for PII and… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/92k-real-world-call-center-scripts-english.10K<n<100K35 likes558 downloads1y agoHugging Face04NoNameFactory /callcentertext10K<n<100K0 likes206 downloads2y agoHugging Face05josemancharo /apptek_callcenter_dialogues_travel_hospitality_no_transcripts AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts) This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case. Changes from the source dataset Restricted the dataset to the travel and hospitality domains. Removed the transcript field (text) entirely. Kept the original audio and the domain, gender, and accent metadata. Preserved the source dataset's test split. This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.audioaudio-classificationn<1K0 likes163 downloads1mo agoHugging Face06MaratDV /russian-call-center-speech-ru 📖 Описание на русском ОписаниеКрупный датасет реальных записей колл-центров на русском языке.Телефонное качество, разговоры «клиент–оператор».Подходит для обучения систем ASR (распознавание речи), NLP, голосовых ассистентов и анализа диалогов. Технические характеристики Язык: русский Общая продолжительность: ~832 часа Формат: MP3 Каналы: моно (клиент и оператор в одном канале) Частота дискретизации: 8000 Гц Битрейт: 32 кбит/с Метаданные: отсутствуют… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/russian-call-center-speech-ru.audioautomatic-speech-recognition1 likes114 downloads1y agoHugging Face07mindweave /call-center-records Call Center Records & Agent Performance Dataset (Free Sample) This is a free sample with 2,013 rows. The full dataset has 19,959 rows across 3 tables. Call detail records for a simulated insurance company call center with 25 agents handling 35,000 calls over 12 months. Includes IVR menu paths, agent assignments, call outcomes, hold times, transfer chains, and customer satisfaction scores. Features realistic patterns: Monday morning surge, lunch dip, seasonal peaks during open… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/call-center-records.tabulartabular-classification1K<n<10K0 likes103 downloads6mo agoHugging Face08hussxamg04 /apptek_callcenter_dialogues AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions. 128.6 hours of speech 14 English accent groups 16 service domains 5–15 minute conversations (long-form) Split-channel audio (one speaker… See the full description on the dataset page: https://huggingface.co/datasets/hussxamg04/apptek_callcenter_dialogues.audioautomatic-speech-recognition1K<n<10K0 likes77 downloads5mo agoHugging Face09urvog /llama2_transcripts_healthcare_callcentertext1K<n<10K5 likes64 downloads3y agoHugging Face10AxonData /multilingual-call-center-speech-dataset Multilingual Call Center Speech Recognition Dataset: 10,000 Hours Dataset Summary 10,000 hours of real-world call center speech recordings in 7 languages with transcripts. Train speech recognition, sentiment analysis, and conversation AI models on authentic customer support audio. Covers support, sales, billing, finance, and pharma domains Dataset Features 📊 Scale & Quality 10,000 hours of inbound & outbound calls Real-world field… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/multilingual-call-center-speech-dataset.audion<1K1 likes59 downloads7mo agoHugging Face11AIxBlock /Human-to-machine-Japanese-audio-call-center-conversations Dataset Card for Japanese audio call center human to machine conversations This dataset contains synthetic audio conversations in Japanese between human customers and machine agents, simulating real-world call center scenarios Dataset Details Dataset Description Curated by: AIxBlock (aixblock.io) Funded by [optional]: AIxBlock (aixblock.io) Shared by [optional]: AIxBlock (aixblock.io) Language(s) (NLP): Japanese License: Creative Commons Attribution Non… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Human-to-machine-Japanese-audio-call-center-conversations.audiotext-to-speechn<1K2 likes58 downloads1y agoHugging Face12minemaster01 /call_center_data_csvtabular10K<n<100K0 likes54 downloads1y agoHugging Face13ud-nlp /call-center-audio Call Center Dataset - 13,000+ Hours Dataset is a large audio dataset containing 13,000+ hours of real-world customer service calls from global call centers, featuring 90%+ unique speakers and time-stamped transcripts for accurate speech recognition, speaker diarization, and conversational AI model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of real customer service calls Data types Audio Tasks Speech… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/call-center-audio.audioautomatic-speech-recognitionn<1K1 likes54 downloads6mo agoHugging Face14AIxBlock /Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains. 🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions. 🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.audioautomatic-speech-recognitionn<1K5 likes51 downloads1y agoHugging Face15UniDataPro /call-center-audio Call Center Dataset The datasets contain over 13,000+ hours of real-world conversations and feature 90%+ unique speakers. It is designed for advancing conversational AI and speech technologies. The dataset provides high-quality, time-stamped transcripts for model training in speech recognition and speaker diarization, enabling businesses to build and refine their AI systems. By utilizing this dataset, researchers and developers can focus on analyzing customer interactions… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/call-center-audio.audiotext-to-speechn<1K1 likes49 downloads1mo agoHugging Face16yevgeniy03 /home-telecom-callcenter-transcripts-viewer Home/Telecom Call Center Transcript Viewer V2 This public dataset is a flattened viewer copy of one source archive from AIxBlock/92k-real-world-call-center-scripts-english: home_ervice_inbound&telecom _outbound.zip It was converted so Hugging Face Data Studio can display the contents as regular Parquet tables. Splits default/train: one row per transcript, all 3,239 transcripts from the source zip. turns/turns: one row per inferred timestamp/sentence turn, linked by… See the full description on the dataset page: https://huggingface.co/datasets/yevgeniy03/home-telecom-callcenter-transcripts-viewer.tabular100K<n<1M0 likes49 downloads5mo agoHugging Face17AxonData /french-call-center-speech-dataset French Call Center Speech Dataset: 1,000+ Hours with Transcripts 1,000+ hours of real-world French call center audio with transcripts. Train speech recognition, sentiment analysis, and customer support AI models on authentic telephone conversations Dataset Summary Key Features ✅ 1,000+ hours of inbound & outbound calls✅ 100% French telephone conversations✅ Real-world audio - no synthetic data✅ Full transcripts in French and in English Full… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/french-call-center-speech-dataset.audion<1K0 likes47 downloads7mo agoHugging Face18InfoBayAI /Arabic-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 64,027 hours of processed Arabic (AR) single-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes43 downloads9d agoHugging Face19InfoBayAI /Hindi_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 1,587,658 hours of processed Hindi dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes42 downloads9d agoHugging Face20InfoBayAI /Nepali_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 229,645 hours of processed Nepalese (NP) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face21InfoBayAI /French_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 31,106 hours of processed French (FR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face22InfoBayAI /Filipino-Tagalog-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4019 hours of processed Tagalog (TL) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino-Tagalog-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face23InfoBayAI /Hindi-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 1,587,658 hours of processed Hindi single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behaviour commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K1 likes40 downloads9d agoHugging Face24InfoBayAI /Arabic_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 64,027 hours of processed Arabic (AR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes39 downloads9d agoHugging Face25InfoBayAI /English-United-Kingdom-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 90,334 hours of processed English (UK) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behavior commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-United-Kingdom-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes39 downloads9d agoHugging Face26InfoBayAI /Russian_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes37 downloads9d agoHugging Face27InfoBayAI /Marathi-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 58,486 hours of processed Marathi (MR) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes36 downloads9d agoHugging Face28InfoBayAI /English_United_States_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 250,362 hours of processed English (US) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_States_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K1 likes35 downloads10d agoHugging Face29InfoBayAI /Odia_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 12,794 hours of processed Odia(Oriya) (OR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Odia_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes35 downloads9d agoHugging Face30InfoBayAI /English_United_Kingdom_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 90,334 hours of processed English (UK) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_Kingdom_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K2 likes34 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.