datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apptek_callcenter_dialogues
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents
across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions.
128.6 hours of speech
14 English accent groups
16 service domains
5–15 minute conversations (long-form)
Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.call-center
Call Center
This is a collection for evaluation of LLMs on call center scenarios.
synthetic_call_center_summaries
Synthetic Call Center Summaries Dataset
Overview
This dataset contains synthetic summaries of call center conversations generated by different prompt configurations.
Each record (in JSON Lines format) includes:
The original dialogue metadata.
A generated summary tailored to provide quick insights for call center service agents.
Extensive evaluation metrics and attributes such as conciseness, formatting, contextual relevance, tone, and actionability.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/synthetic_call_center_summaries.isha-call-center-qa-datavi-callcenter-silver-evalcall-center-qa-squad-2
