CoolFace
Datasetpublic

aranemini/hawrami-kurdish-raw-audio

Hawrami Raw Audio Collection Overview This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources. The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content. The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/hawrami-kurdish-raw-audio.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes125downloads
Dataset Card

Hawrami Raw Audio Collection

Overview

This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources.

The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content.

The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other naturally occurring spoken interactions. The dataset captures a variety of speakers, speaking styles, and recording conditions representative of real-world Hawrami speech.

The dataset is intended to support research on:

  • —Automatic Speech Recognition (ASR)
  • —Speech Translation (ST)
  • —Text-to-Speech (TTS)
  • —Self-supervised Learning (SSL)

Important Notice

This dataset was collected for research and educational purposes.

If you are the owner of any content included in this repository and would like it removed, please contact:

emini.aran@gmail.com

Dataset Statistics

StatisticValue
LanguageHawrami Kurdish
Duration~500 hours
Data TypeRaw audio
Main SourceRocyar (Sterk TV)
Speech StylePredominantly spontaneous and semi-spontaneous speech

Citation

bibtex
@inproceedings{mohammadamini25_interspeech,
  title={Scaling pseudo-labeling data for end-to-end low-resource speech translation (the case of Kurdish language)},
  author={Mohammad Mohammadamini and Aghilas Sini and Marie Tahon and Antoine Laurent},
  booktitle={Interspeech 2025},
  year={2025}
}

Disclaimer

The audio recordings remain subject to the rights of their original creators and publishers. This repository does not claim ownership of the underlying content and is provided solely for research and educational use.

Although the dataset was collected to represent Hawrami Kurdish speech, users should expect some variation in speakers, accents, and speaking styles across the collection.