datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
northern-kurdish-raw-audio
Northern Kurdish Raw Audio Collection
Overview
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The collection was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-supervised Learning (SSL)
Spoken Language Understanding (SLU)
The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.northern-kurdish-pseudolabel
Northern Kurdish Raw Audio Collection
Dataset Summary
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The corpus was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-Supervised Learning (SSL)
Spoken Language Understanding (SLU)
Low-Resource Speech Processing
The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.northeastbench-speech
NortheastBench-Speech
A standardized 8,099-utterance evaluation suite for automatic speech recognition and speech-text retrieval across eight indigenous and regional languages of Northeast India, spanning three language families: Austroasiatic (Khasi), Tibeto-Burman (Garo, Mizo, Kokborok, Wancho, Chakma), and Indo-Aryan (Nagamese, Assamese).
Introduced in NE-MultiSpeech: A Multilingual Speech Corpus and ASR Benchmark for Northeast Indian Languages.
Languages and… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeastbench-speech.lj_real_estate_deposition_full_case
LJ Real Estate Deposition – Full Case (Anonymized Audio + Transcript, EN)
Overview
This dataset contains a full real-world legal deposition in a real-estate / business context, provided as anonymized long-form audio + full court-style transcript.
It is designed for teams building and evaluating:
Automatic speech recognition (ASR) for long-form legal speech
Legal / real-estate conversation and dialogue models
Agent-style systems that need realistic, high-stakes… See the full description on the dataset page: https://huggingface.co/datasets/NorthAlabamaConsultants/lj_real_estate_deposition_full_case.
