datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
malayalam-asr-corpus
Malayalam ASR Corpus
A multi-corpus Malayalam speech dataset aggregated from 5 public sources,
created for fine-tuning ASR models on Malayalam language.
This dataset was used to train
sajilck/whisper-small-malayalam
— the first multi-corpus Malayalam Whisper model on HuggingFace,
achieving 37.64% WER on CommonVoice 25 Malayalam test set.
Source Corpora
Corpus
Domain
Speaker Type
License
IMaSC
TTS / Read speech
Studio speakers
CC BY 4.0
SMC Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/sajilck/malayalam-asr-corpus.malayalam-asr-5K
Malayalam ASR 5K — Verified Anchor Set
5,225 manually verified Malayalam speech-transcript pairs (7.1 hours),
speaker-disjoint across train/dev/test. Every record in this release
carries is_verified: true — each transcript was checked, not machine-
generated and left unreviewed.
Splits (speaker-disjoint, source-aware)
split
utterances
%
disjoint units
speakers
hours
train
3,657
70%
9
7
4.8
dev
783
15%
5
5
1.1
test
785
15%
6
5
1.1
No speaker or… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-asr-5K.malayalam_speech
Malayalam Speech Dataset (VibeVoice)
This dataset contains Malayalam speech audio files and their corresponding transcriptions, prepared for fine-tuning ASR (Automatic Speech Recognition) models like VibeVoice.
Dataset Description
The dataset consists of approximately 4,126 Malayalam audio clips, split into training and testing sets with a 90/10 ratio, stratified by speaker gender.
Total Audio Files: 4,126
Train Split: 3,712 samples
Test Split: 414 samples
Total… See the full description on the dataset page: https://huggingface.co/datasets/ArjunJ/malayalam_speech.vividh-test-malayalam
🎙️ Vividh-ASR Benchmark — Malayalam (Test Split)
How well does your ASR model actually work in the wild?Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart.
Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-malayalam.malayalam-tts-pro-voice
Malayalam Speech Dataset (Text + Audio)
This dataset contains Malayalam speech audio clips paired with text transcripts.It is designed for training and fine-tuning ASR (Automatic Speech Recognition),TTS (Text-to-Speech) models, and speech-to-speech translation systems.
Emotions Added
giggles
laughs
long pause
chuckles
whispers
gasps
clears throat
singing
laughs nervously
burps
exhales
📁 Dataset Structure
Column
Description
audio
Path to the… See the full description on the dataset page: https://huggingface.co/datasets/sachin6624/malayalam-tts-pro-voice.malayalam-whisper-corpus-v2
Malayalam Whisper Corpus v2
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus-v2.Malayalam-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 14,980 hours of processed Malayalam (ML) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Call-Center-Audio-Dataset-Single-Channel.Malayalam_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 14,980 hours of processed Malayalam (ML) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Call_Center_Audio_Dataset_Dual_Channel.Malayalam_Podcast_Audio_DatasetDataset Description:
This dataset is a large-scale collection of 3,956 hours of processed Malayalam podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world interactions across diverse topics and formats. The dataset preserves natural speech patterns, speaker variability, and authentic podcast environments, making it… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Podcast_Audio_Dataset.malayalam-whisper-corpus_v3
Malayalam Whisper Corpus v3
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus_v3.malayalam-tts-pro-voice
Malayalam Speech Dataset (Text + Audio)
This dataset contains Malayalam speech audio clips paired with text transcripts.It is designed for training and fine-tuning ASR (Automatic Speech Recognition),TTS (Text-to-Speech) models, and speech-to-speech translation systems.
Emotions Added
giggles
laughs
long pause
chuckles
whispers
gasps
clears throat
singing
laughs nervously
burps
exhales
📁 Dataset Structure
Column
Description
audio… See the full description on the dataset page: https://huggingface.co/datasets/trysem/malayalam-tts-pro-voice.Malayalam_Podcast_Audio_Dataset_Dual_Channel
Dataset Description
This dataset is a large-scale collection of 3,956 hours of processed Malayalam dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Podcast_Audio_Dataset_Dual_Channel.Malayalam-Speech-Dataset
🎧 Malayalam Speech Dataset
The Malayalam Speech Dataset is a high-quality speech audio dataset designed to power AI and machine learning systems with reliable and diverse audio data. It includes 95 hours of recorded speech data across 650 files, available in MP3 and WAV formats, with a total size of 227 MB. This well-structured audio dataset delivers balanced and representative voice data, featuring 54% female and 46% male speakers, with age groups ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Malayalam-Speech-Dataset.
