call center
apptek_callcenter_dialogues
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents
across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions.
128.6 hours of speech
14 English accent groups
16 service domains
5–15 minute conversations (long-form)
Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.call-center-prod-data92k-real-world-call-center-scripts-englishArXiv Paper Publication Here: "Real-World En Call Center Transcripts Dataset with PII Redaction"
This dataset includes 91,706 high-quality transcriptions corresponding to approximately 10,500 hours of real-world call center conversations in English, collected across various industries and global regions. The dataset features both inbound and outbound calls and spans multiple accents, including Indian, American, and Filipino English. All transcripts have been carefully redacted for PII and… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/92k-real-world-call-center-scripts-english.callcenterapptek_callcenter_dialogues_travel_hospitality_no_transcripts
AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts)
This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case.
Changes from the source dataset
Restricted the dataset to the travel and hospitality domains.
Removed the transcript field (text) entirely.
Kept the original audio and the domain, gender, and accent metadata.
Preserved the source dataset's test split.
This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.russian-call-center-speech-ru
📖 Описание на русском
ОписаниеКрупный датасет реальных записей колл-центров на русском языке.Телефонное качество, разговоры «клиент–оператор».Подходит для обучения систем ASR (распознавание речи), NLP, голосовых ассистентов и анализа диалогов.
Технические характеристики
Язык: русский
Общая продолжительность: ~832 часа
Формат: MP3
Каналы: моно (клиент и оператор в одном канале)
Частота дискретизации: 8000 Гц
Битрейт: 32 кбит/с
Метаданные: отсутствуют… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/russian-call-center-speech-ru.
