aranemini/southern-kurdish-raw-audio
Southern Kurdish Raw Audio Collection Overview This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs. The main sources are Aryen TV and Kurd Channel. The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-raw-audio.
Southern Kurdish Raw Audio Collection
Overview
This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs.
The main sources are Aryen TV and Kurd Channel.
The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is Southern Kurdish, some recordings—particularly interviews and discussion programs—may contain contributions from speakers of other Kurdish varieties, especially Central Kurdish (Sorani).
The dataset is intended to support research on:
- Automatic Speech Recognition (ASR)
- Speech Translation (ST)
- Text-to-Speech (TTS)
- Self-supervised Learning (SSL)
Important Notice
This dataset was collected for research and educational purposes.
If you are the owner of any content included in this repository and would like it removed, please contact:
emini.aran@gmail.com
Dataset Statistics
Citation
@inproceedings{mohammadamini25_interspeech,
title={Scaling pseudo-labeling data for end-to-end low-resource speech translation (the case of Kurdish language)},
author={Mohammad Mohammadamini and Aghilas Sini and Marie Tahon and Antoine Laurent},
booktitle={Interspeech 2025},
year={2025}
}Disclaimer
The audio recordings remain subject to the rights of their original creators and publishers. This repository does not claim ownership of the underlying content and is provided solely for research and educational use.
Although the dataset was collected to represent Southern Kurdish speech, users should expect a limited amount of dialectal variation due to the presence of guest speakers and interviews from other Kurdish-speaking regions.
