datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arabic_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 64,027 hours of processed Arabic (AR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic_Call_Center_Audio_Dataset_Dual_Channel.Arabic-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 64,027 hours of processed Arabic (AR) single-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic-Call-Center-Audio-Dataset-Single-Channel.Hindi_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 1,587,658 hours of processed Hindi dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi_Call_Center_Audio_Dataset_Dual_Channel.Russian_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian_Call_Center_Audio_Dataset_Dual_Channel.Nepali_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 229,645 hours of processed Nepalese (NP) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel.French_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 31,106 hours of processed French (FR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French_Call_Center_Audio_Dataset_Dual_Channel.Filipino-Tagalog-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4019 hours of processed Tagalog (TL) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino-Tagalog-Call-Center-Audio-Dataset-Single-Channel.Hindi-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 1,587,658 hours of processed Hindi single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behaviour commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Call-Center-Audio-Dataset-Single-Channel.English_United_States_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 250,362 hours of processed English (US) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_States_Call_Center_Audio_Dataset_Dual_Channel.English-United-Kingdom-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 90,334 hours of processed English (UK) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, interruptions, and natural speaking behavior commonly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-United-Kingdom-Call-Center-Audio-Dataset-Single-Channel.English_United_Kingdom_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 90,334 hours of processed English (UK) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_United_Kingdom_Call_Center_Audio_Dataset_Dual_Channel.Filipino_Tagalog_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4,019 hours of processed Tagalog (TL) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino_Tagalog_Call_Center_Audio_Dataset_Dual_Channel.Marathi-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 58,486 hours of processed Marathi (MR) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-Call-Center-Audio-Dataset-Single-Channel.German_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 212 hours of processed German (DE) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/German_Call_Center_Audio_Dataset_Dual_Channel.Odia_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 12,794 hours of processed Odia(Oriya) (OR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Odia_Call_Center_Audio_Dataset_Dual_Channel.Mizo_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 468 hours of processed Mizo (MZ) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Mizo_Call_Center_Audio_Dataset_Dual_Channel.English_India_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 30,320 processed English (India) dual-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments.… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_India_Call_Center_Audio_Dataset_Dual_Channel.Malayalam-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 14,980 hours of processed Malayalam (ML) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Call-Center-Audio-Dataset-Single-Channel.Russian-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian-Call-Center-Audio-Dataset-Single-Channel.Bengali_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali_Call_Center_Audio_Dataset_Dual_Channel.Nepali_Call_Center_Audio_Dataset_Dual_Channel
Nepali Call Center Audio Dataset — Dual Channel
Source and Attribution
This repository contains data obtained from the original dataset
published by InfoBay AI Ltd.
Original Dataset
Name:
Nepali_Call_Center_Audio_Dataset_Dual_Channel
Original repository:
https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel
Original creator:
InfoBay AI Ltd.
The original dataset is listed as CC BY 4.0 on its Hugging Face
repository.… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose7777/Nepali_Call_Center_Audio_Dataset_Dual_Channel.Spanish_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 452 hours of processed Spanish (MX) and 785 hours of processed Spanish (ES) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Spanish_Call_Center_Audio_Dataset_Dual_Channel.Somali-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 952 hours of processed Somali (SO) and 105 hours of processed Somali (UG) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Somali-Call-Center-Audio-Dataset-Single-Channel.Swahili_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 194,331 hours of processed Swahili (SW) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Swahili_Call_Center_Audio_Dataset_Dual_Channel.Somali_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 952 hours of processed Somali (SO) and 105 hours of processed Somali (UG) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Somali_Call_Center_Audio_Dataset_Dual_Channel.Tamil_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 15,056 hours of processed Tamil (TA) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Tamil_Call_Center_Audio_Dataset_Dual_Channel.Bengali-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali-Call-Center-Audio-Dataset-Single-Channel.English-United-States-Call-Center-Audio-Dataset-Single-Channel
Dataset Description
This dataset is a large-scale collection of 250,362 hours of processed English (US) single-channel call center audio recordings, part of a broader multilingual dataset spanning 3,569,083 hours across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
This dataset is particularly valuable for building scalable enterprise-grade AI systems including Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-United-States-Call-Center-Audio-Dataset-Single-Channel.French-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 31,106 hours of processed French (FR) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French-Call-Center-Audio-Dataset-Single-Channel.Gujarati_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 23,361 hours of processed Gujarati (GU) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Gujarati_Call_Center_Audio_Dataset_Dual_Channel.
