datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hawrami_speech
Hawrami Speech (hawrami_speech)
Studio-recorded Hawrami (Hewramî) read speech with sentence-level transcriptions:
4,977 utterances over ~4 hours of audio, with speaker_id labels covering 102
distinct speakers.
At a glance
Rows
4,977 — train 4,773 / test 204
Columns
audio, sentence, gender, language, original_full_path, duration, speaker_id
Parquet on disk
467.0 MB
Audio format
WAV (files like voice__24773.wav)
Language
Hawrami (language is… See the full description on the dataset page: https://huggingface.co/datasets/razhan/hawrami_speech.malayalam-whisper-corpus-v2
Malayalam Whisper Corpus v2
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus-v2.malayalam-whisper-corpus_v3
Malayalam Whisper Corpus v3
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus_v3.
