datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VCTKmalagasy-speech-fullanv_data_ke_kikuyu_mergedanv_data_ke_kikuyu_scriptedkinyarwanda-speech-500h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.anv-data-ke-somali-fullethiopian-speech-flat
Ethio Speech Copus — Afrivoices Ethiopian
📌 Overview
The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages.
It is designed to support the development of speech-to-text technologies for low-resource languages.
This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages.
The broader goal of the initiative is to collect 600 hours… See the full description on the dataset page: https://huggingface.co/datasets/badrex/ethiopian-speech-flat.kinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.swahili-speech-400hrshona-speechLJSpeech-1.1
The LJ Speech Dataset
Version 1.1
July 5, 2017
https://keithito.com/LJ-Speech-Dataset
OVERVIEW
This is a public domain speech dataset consisting of 13,100 short audio clips
of a single speaker reading passages from 7 non-fiction books. A transcription
is provided for each clip. Clips vary in length from 1 to 10 seconds and have
a total length of approximately 24 hours.
The texts were published between 1884 and 1964, and are in the public domain.
The audio was recorded in… See the full description on the dataset page: https://huggingface.co/datasets/badayvedat/LJSpeech-1.1.wolaytta-speechfleurs-badini
FLEURS-Badini
Dataset Summary
FLEURS-Badini is a speech dataset for the Badini dialect of Northern Kurdish, designed for research in:
Automatic Speech Recognition (ASR)
Speech-to-Text Translation (S2TT)
It is a dialect-specific extension of the FLEURS benchmark, providing aligned speech–text–translation data for a low-resource language variant.
The dataset contains 5,224 utterances (~15h40m) recorded from 45 speakers.
Supported Tasks
Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/BadiniSpeechNLP/fleurs-badini.badini-asr-benchmark
Dataset Card for "badini-asr-benchmark"
More Information needed
amharic-speechMADIS5-spoken-arabic-dialects
Dataset Overview
MADIS-5 (Multi-domain Arabic Dialect Identification in Speech) is a manually curated dataset
designed to facilitate evaluation of cross-domain robustness of Arabic Dialect Identification (ADI) systems.
This dataset provides a comprehensive benchmark for testing out-of-domain generalization across different speech domains with diverse recording conditions and speaking styles.
Dataset Statistics
Total Duration: ~12 hours of speech
Total… See the full description on the dataset page: https://huggingface.co/datasets/badrex/MADIS5-spoken-arabic-dialects.kalenjin-speech-fullanv-data-ke-somaliarabic-speech-SADA22-MSA
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.afaan_oromo-speechfulani-speechethiopian-speechbadini-tts-samplestigrinya-speecharabic-speech-SADA22-Khaliji
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.waxalNLP-ethiopic-finalBA_Datensatz_V2Bad_Therapeutic_Musicanv-data-ke-somali-testsidama-speech
