datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hawrami-kurdish-raw-audio
Hawrami Raw Audio Collection
Overview
This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources.
The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content.
The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/hawrami-kurdish-raw-audio.hawrami_speech
Hawrami Speech (hawrami_speech)
Studio-recorded Hawrami (Hewramî) read speech with sentence-level transcriptions:
4,977 utterances over ~4 hours of audio, with speaker_id labels covering 102
distinct speakers.
At a glance
Rows
4,977 — train 4,773 / test 204
Columns
audio, sentence, gender, language, original_full_path, duration, speaker_id
Parquet on disk
467.0 MB
Audio format
WAV (files like voice__24773.wav)
Language
Hawrami (language is… See the full description on the dataset page: https://huggingface.co/datasets/razhan/hawrami_speech.malayalam-whisper-corpus-v2
Malayalam Whisper Corpus v2
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus-v2.malayalam-whisper-corpus_v3
Malayalam Whisper Corpus v3
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus_v3.
