datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
central-kurdish-pseudolabel
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus
Dataset Summary
This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).
The dataset was automatically generated using a pipeline composed of:
Speech segmentation
Automatic Speech Recognition (ASR)
Machine Translation (MT)
The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.northern-kurdish-pseudolabel
Northern Kurdish Raw Audio Collection
Dataset Summary
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The corpus was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-Supervised Learning (SSL)
Spoken Language Understanding (SLU)
Low-Resource Speech Processing
The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.central-kurdish-tts4all
TTS4All Central Kurdish Speech Dataset
Dataset Summary
The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish).
The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers.
The corpus was designed to support:
Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.kurdish-voice-dataset
🎙️ داتاسێتی دەنگی ڕەنج سەنگاوی (Kurdish Voice Dataset - Ranj Sangawi)
١,٢٠٠,٠٠٠+ وشەی پوخت | ١٠٠,٠٠٠ کلیپ | ١٦٦.٠ کاتژمێر دەنگ
گەورەترین، دەوڵەمەندترین و پرۆفێشناڵترین داتاسێتی دەنگی کوردیی سۆرانی لەسەر ئاستی جیهان، تایبەتکراو ١٠٠٪ بۆ بێژەر و پێشکەشکار ڕەنج سەنگاوی (Ranj Sangawi) بەرهەمهێنراو لەلایەن فەرمان عوسمان (Farman Othman).
ئەم داتاسێتە کۆکراوەی وتەکان، دیبەیتەکان، دەقە ڕۆژنامەوانییەکان، وتووێژە تەلەفزیۆنییەکان (لە بەرنامەی «لەگەڵ ڕەنج» لە کەناڵی ڕووداو،… See the full description on the dataset page: https://huggingface.co/datasets/farmanOthman/kurdish-voice-dataset.southern-kurdish-asr
Bestun: Southern Kurdish speech recognition resources and benchmarking
This repository provides speech recognition (ASR) resources for Southern Kurdish (ISO 639-3: sdh), a threatened variant of the Kurdish macrolanguage.It includes:
Bestun training corpus: ~30 hours of manually validated read speech
Evaluation benchmark: 773 validated utterances (86.74 minutes) recorded by 8 speakers from different Southern Kurdish vernacular regions
The dataset and models are released under CC… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-asr.kurdish-multidialect-asr-benchmark
Kurdish Dialect Speech Corpus
This project aims to provide a multi-dialect speech recognition benchmark for the Kurdish language. The Central Kurdish portion is the same as the Asosoft benchmark. The sentences were originally written in Central Kurdish (CKB), translated into other Kurdish dialects, and then recorded by native speakers.
The current version includes three Kurdish dialects: Central Kurdish, Northern Kurdish, and Southern Kurdish. A Hawrami version and the Badini… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/kurdish-multidialect-asr-benchmark.
