CoolFace
Datasetpublicgated

InfoBayAI/Malayalam_Podcast_Audio_Dataset

Dataset Description: This dataset is a large-scale collection of 3,956 hours of processed Malayalam podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It captures real-world interactions across diverse topics and formats. The dataset preserves natural speech patterns, speaker variability, and authentic podcast environments, making it… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Podcast_Audio_Dataset.

sourceHugging Facecc-by-4.0updated 10d agoView on Hugging Face
0likes26downloads
settings

This repository belongs to InfoBayAI on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameMalayalam_Podcast_Audio_Dataset
visibilitypublic
licencecc-by-4.0
gatedyes
ownerInfoBayAI
Account settings
InfoBayAI/Malayalam_Podcast_Audio_Dataset · CoolFace