CoolFace
Datasetpublic

aranemini/southern-kurdish-raw-audio

Southern Kurdish Raw Audio Collection Overview This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs. The main sources are Aryen TV and Kurd Channel. The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-raw-audio.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes196downloads
Dataset Card

Southern Kurdish Raw Audio Collection

Overview

This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs.

The main sources are Aryen TV and Kurd Channel.

The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is Southern Kurdish, some recordings—particularly interviews and discussion programs—may contain contributions from speakers of other Kurdish varieties, especially Central Kurdish (Sorani).

The dataset is intended to support research on:

  • Automatic Speech Recognition (ASR)
  • Speech Translation (ST)
  • Text-to-Speech (TTS)
  • Self-supervised Learning (SSL)

Important Notice

This dataset was collected for research and educational purposes.

If you are the owner of any content included in this repository and would like it removed, please contact:

emini.aran@gmail.com

Dataset Statistics

StatisticValue
LanguageSouthern Kurdish (SDH)
Duration~170 hours
Data TypeRaw audio
Main SourcesAryen TV, Kurd Channel
Speech StylePredominantly spontaneous and semi-spontaneous speech

Citation

bibtex
@inproceedings{mohammadamini25_interspeech,
  title={Scaling pseudo-labeling data for end-to-end low-resource speech translation (the case of Kurdish language)},
  author={Mohammad Mohammadamini and Aghilas Sini and Marie Tahon and Antoine Laurent},
  booktitle={Interspeech 2025},
  year={2025}
}

Disclaimer

The audio recordings remain subject to the rights of their original creators and publishers. This repository does not claim ownership of the underlying content and is provided solely for research and educational use.

Although the dataset was collected to represent Southern Kurdish speech, users should expect a limited amount of dialectal variation due to the presence of guest speakers and interviews from other Kurdish-speaking regions.