CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sajilck /malayalam-asr-corpus Malayalam ASR Corpus A multi-corpus Malayalam speech dataset aggregated from 5 public sources, created for fine-tuning ASR models on Malayalam language. This dataset was used to train sajilck/whisper-small-malayalam — the first multi-corpus Malayalam Whisper model on HuggingFace, achieving 37.64% WER on CommonVoice 25 Malayalam test set. Source Corpora Corpus Domain Speaker Type License IMaSC TTS / Read speech Studio speakers CC BY 4.0 SMC Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/sajilck/malayalam-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes316 downloads2mo agoHugging Face02icfoss /malayalam-asr-5K Malayalam ASR 5K — Verified Anchor Set 5,225 manually verified Malayalam speech-transcript pairs (7.1 hours), speaker-disjoint across train/dev/test. Every record in this release carries is_verified: true — each transcript was checked, not machine- generated and left unreviewed. Splits (speaker-disjoint, source-aware) split utterances % disjoint units speakers hours train 3,657 70% 9 7 4.8 dev 783 15% 5 5 1.1 test 785 15% 6 5 1.1 No speaker or… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-asr-5K.audioautomatic-speech-recognition1K<n<10K0 likes94 downloads11d agoHugging Face03ArjunJ /malayalam_speech Malayalam Speech Dataset (VibeVoice) This dataset contains Malayalam speech audio files and their corresponding transcriptions, prepared for fine-tuning ASR (Automatic Speech Recognition) models like VibeVoice. Dataset Description The dataset consists of approximately 4,126 Malayalam audio clips, split into training and testing sets with a 90/10 ratio, stratified by speaker gender. Total Audio Files: 4,126 Train Split: 3,712 samples Test Split: 414 samples Total… See the full description on the dataset page: https://huggingface.co/datasets/ArjunJ/malayalam_speech.audioautomatic-speech-recognition1K<n<10K0 likes93 downloads8mo agoHugging Face04adalat-ai /vividh-test-malayalam 🎙️ Vividh-ASR Benchmark — Malayalam (Test Split) How well does your ASR model actually work in the wild?Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart. Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-malayalam.audioautomatic-speech-recognition10K<n<100K1 likes45 downloads4mo agoHugging Face05sachin6624 /malayalam-tts-pro-voice Malayalam Speech Dataset (Text + Audio) This dataset contains Malayalam speech audio clips paired with text transcripts.It is designed for training and fine-tuning ASR (Automatic Speech Recognition),TTS (Text-to-Speech) models, and speech-to-speech translation systems. Emotions Added giggles laughs long pause chuckles whispers gasps clears throat singing laughs nervously burps exhales 📁 Dataset Structure Column Description audio Path to the… See the full description on the dataset page: https://huggingface.co/datasets/sachin6624/malayalam-tts-pro-voice.audiotext-to-speechn<1K3 likes41 downloads9mo agoHugging Face06hawks23 /malayalam-whisper-corpus-v2 Malayalam Whisper Corpus v2 Dataset Description A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper. Sources Mozilla Common Voice 11.0 (CC0-1.0) OpenSLR Malayalam Speech Corpus (Apache 2.0) Indic Speech 2022 Challenge (CC-BY-4.0) IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus-v2.audioautomatic-speech-recognition10K<n<100K0 likes35 downloads1y agoHugging Face07InfoBayAI /Malayalam-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 14,980 hours of processed Malayalam (ML) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes32 downloads10d agoHugging Face08InfoBayAI /Malayalam_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 14,980 hours of processed Malayalam (ML) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes26 downloads10d agoHugging Face09InfoBayAI /Malayalam_Podcast_Audio_DatasetgatedDataset Description: This dataset is a large-scale collection of 3,956 hours of processed Malayalam podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It captures real-world interactions across diverse topics and formats. The dataset preserves natural speech patterns, speaker variability, and authentic podcast environments, making it… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Podcast_Audio_Dataset.audioautomatic-speech-recognitionn<1K0 likes24 downloads9d agoHugging Face10hawks23 /malayalam-whisper-corpus_v3 Malayalam Whisper Corpus v3 Dataset Description A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper. Sources Mozilla Common Voice 11.0 (CC0-1.0) OpenSLR Malayalam Speech Corpus (Apache 2.0) Indic Speech 2022 Challenge (CC-BY-4.0) IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus_v3.audioautomatic-speech-recognition10K<n<100K0 likes20 downloads1y agoHugging Face11trysem /malayalam-tts-pro-voice Malayalam Speech Dataset (Text + Audio) This dataset contains Malayalam speech audio clips paired with text transcripts.It is designed for training and fine-tuning ASR (Automatic Speech Recognition),TTS (Text-to-Speech) models, and speech-to-speech translation systems. Emotions Added giggles laughs long pause chuckles whispers gasps clears throat singing laughs nervously burps exhales 📁 Dataset Structure Column Description audio… See the full description on the dataset page: https://huggingface.co/datasets/trysem/malayalam-tts-pro-voice.audiotext-to-speechn<1K0 likes17 downloads4mo agoHugging Face12InfoBayAI /Malayalam_Podcast_Audio_Dataset_Dual_Channelgated Dataset Description This dataset is a large-scale collection of 3,956 hours of processed Malayalam dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam_Podcast_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes16 downloads10d agoHugging Face13Speech-data /Malayalam-Speech-Dataset 🎧 Malayalam Speech Dataset The Malayalam Speech Dataset is a high-quality speech audio dataset designed to power AI and machine learning systems with reliable and diverse audio data. It includes 95 hours of recorded speech data across 650 files, available in MP3 and WAV formats, with a total size of 227 MB. This well-structured audio dataset delivers balanced and representative voice data, featuring 54% female and 46% male speakers, with age groups ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Malayalam-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes15 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.