CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aranemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.2k downloads3mo agoHugging Face02aranemini /northern-kurdish-raw-audio Northern Kurdish Raw Audio Collection Overview This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources. The collection was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-supervised Learning (SSL) Spoken Language Understanding (SLU) The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.audioautomatic-speech-recognition1K<n<10K0 likes754 downloads3mo agoHugging Face03aranemini /northern-kurdish-pseudolabel Northern Kurdish Raw Audio Collection Dataset Summary This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources. The corpus was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-Supervised Learning (SSL) Spoken Language Understanding (SLU) Low-Resource Speech Processing The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.audioautomatic-speech-recognition100K<n<1M1 likes223 downloads3mo agoHugging Face04aranemini /southern-kurdish-raw-audio Southern Kurdish Raw Audio Collection Overview This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs. The main sources are Aryen TV and Kurd Channel. The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-raw-audio.audioautomatic-speech-recognitionn<1K0 likes197 downloads3mo agoHugging Face05aranemini /central-kurdish-audiobook-raw Central Kurdish Audiobook Raw Audio Collection Overview This repository contains a large collection of raw Central Kurdish (Sorani Kurdish) audiobook recordings gathered from publicly available online sources. The collection was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-supervised learning The dataset contains approximately 4,300 hours of speech collected from 1026… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-audiobook-raw.audioautomatic-speech-recognition10K<n<100K1 likes186 downloads3mo agoHugging Face06aranemini /central-kurdish-tts4all TTS4All Central Kurdish Speech Dataset Dataset Summary The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish). The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers. The corpus was designed to support: Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.audiotext-to-speech10K<n<100K3 likes150 downloads3mo agoHugging Face07farmanOthman /kurdish-voice-dataset 🎙️ داتاسێتی دەنگی ڕەنج سەنگاوی (Kurdish Voice Dataset - Ranj Sangawi) ١,٢٠٠,٠٠٠+ وشەی پوخت | ١٠٠,٠٠٠ کلیپ | ١٦٦.٠ کاتژمێر دەنگ گەورەترین، دەوڵەمەندترین و پرۆفێشناڵترین داتاسێتی دەنگی کوردیی سۆرانی لەسەر ئاستی جیهان، تایبەتکراو ١٠٠٪ بۆ بێژەر و پێشکەشکار ڕەنج سەنگاوی (Ranj Sangawi) بەرهەمهێنراو لەلایەن فەرمان عوسمان (Farman Othman). ئەم داتاسێتە کۆکراوەی وتەکان، دیبەیتەکان، دەقە ڕۆژنامەوانییەکان، وتووێژە تەلەفزیۆنییەکان (لە بەرنامەی «لەگەڵ ڕەنج» لە کەناڵی ڕووداو،… See the full description on the dataset page: https://huggingface.co/datasets/farmanOthman/kurdish-voice-dataset.audiotext-to-speech100K<n<1M0 likes129 downloads3d agoHugging Face08aranemini /hawrami-kurdish-raw-audio Hawrami Raw Audio Collection Overview This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources. The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content. The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/hawrami-kurdish-raw-audio.audioautomatic-speech-recognitionn<1K0 likes125 downloads3mo agoHugging Face09aranemini /southern-kurdish-asr Bestun: Southern Kurdish speech recognition resources and benchmarking This repository provides speech recognition (ASR) resources for Southern Kurdish (ISO 639-3: sdh), a threatened variant of the Kurdish macrolanguage.It includes: Bestun training corpus: ~30 hours of manually validated read speech Evaluation benchmark: 773 validated utterances (86.74 minutes) recorded by 8 speakers from different Southern Kurdish vernacular regions The dataset and models are released under CC… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-asr.audioautomatic-speech-recognition10K<n<100K1 likes86 downloads7mo agoHugging Face10aranemini /kurdish-multidialect-asr-benchmark Kurdish Dialect Speech Corpus This project aims to provide a multi-dialect speech recognition benchmark for the Kurdish language. The Central Kurdish portion is the same as the Asosoft benchmark. The sentences were originally written in Central Kurdish (CKB), translated into other Kurdish dialects, and then recorded by native speakers. The current version includes three Kurdish dialects: Central Kurdish, Northern Kurdish, and Southern Kurdish. A Hawrami version and the Badini… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/kurdish-multidialect-asr-benchmark.audioautomatic-speech-recognition1K<n<10K0 likes67 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.