datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
northern-kurdish-pseudolabel
Northern Kurdish Raw Audio Collection
Dataset Summary
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The corpus was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-Supervised Learning (SSL)
Spoken Language Understanding (SLU)
Low-Resource Speech Processing
The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.northeastbench-speech
NortheastBench-Speech
A standardized 8,099-utterance evaluation suite for automatic speech recognition and speech-text retrieval across eight indigenous and regional languages of Northeast India, spanning three language families: Austroasiatic (Khasi), Tibeto-Burman (Garo, Mizo, Kokborok, Wancho, Chakma), and Indo-Aryan (Nagamese, Assamese).
Introduced in NE-MultiSpeech: A Multilingual Speech Corpus and ASR Benchmark for Northeast Indian Languages.
Languages and… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeastbench-speech.
