datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.global-news-radio-30s
Global News Radio Dataset
Multilingual news radio recordings from 51 languages across 42 countries.
Recordings
51
Total audio
1500 min (25.0 h)
Format
MP3 16kHz mono 64kbps
Parquet shards
11
Languages
51
Countries
42
Size
687 MB
Languages
Amharic, Arabic, Bashkir, Basque, Belarusian, Bengali, Brazilian Portuguese,Portugues Do Brasil,Português Brasil, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, Flemish… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-30s.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/news_youtube_uzbek_speech_dataset.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/news_youtube_uzbek_speech_dataset.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/news_youtube_uzbek_speech_dataset.global-news-radio-debug
Global News Radio Dataset (1 hour per station)
Every news radio station from the Radio Browser API, recorded for 1 hour each.
Attempted
3037
Successful
2553
Failed
484
Total audio
21 hours
Parquet shards
256
Size
0.6 GB
Format
MP3 16kHz mono 64kbps
Usage
from datasets import load_dataset
ds = load_dataset("NathanRoll/global-news-radio-debug", streaming=True)
for sample in ds["train"]:
print(sample["station_name"], sample["language"]… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-debug.ALFFA-Swahili-News
ALFFA Swahili News
Dataset Description
The ALFFA Swahili News dataset is a speech corpus designed for automatic speech recognition (ASR) research in Swahili, an under-resourced African language. This dataset is part of the ALFFA (African Languages in the Field: speech Fundamentals and Automation) project and contains approximately 11.8 hours of broadcast news audio from Radio France International's Swahili service, recorded between November 2010 and March 2011.
The… See the full description on the dataset page: https://huggingface.co/datasets/nickdee96/ALFFA-Swahili-News.The_Arabic_News_speech_Corpus_Dataset
Arabic News Speech Corpus Dataset
This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics.
Dataset Details
Dataset Description
This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.khit_thit_news_voices
Khit Thit News Voices
In the fight for truth, these are the voices that refuse to be silenced.
Khit Thit News Voices is a focused collection of 15,841 audio segments (≈14.7 hours total) from Khit Thit News, one of Myanmar's most vital and trusted independent media outlets. Founded by renowned journalist Mr. Thar Lun Zaung Htet, Khit Thit News stands as a pillar of reliable information and a primary voice for democratic forces within the country.
This dataset primarily features the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/khit_thit_news_voices.mrtv_news_voices
🗣️ Overview
MRTV Voices is a large-scale Burmese speech dataset built from publicly available news broadcasts and programs aired on Myanma Radio and Television (MRTV) — the official state-run media channel of Myanmar.
🎙️ It contains over 130,000 short audio clips (≈117 hours) with aligned transcripts derived from auto-generated subtitles.
This dataset captures:
Formal Burmese used in government bulletins and official reports
Clear pronunciation, enunciation, and pacing —… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mrtv_news_voices.New-Lisan-Sudanese-TTS-Dataset
Lisan Sudanese TTS Dataset
A synthetic Text-to-Speech (TTS) and Automatic Speech Recognition
(ASR) dataset specifically for Sudanese Arabic.
1,878 high-quality sentences featuring 20 synthetic speakers (10
male, 10 female).
Reconstructed from the Lisan-Sudanese Morphological Dataset
(52K manually annotated social media tokens from Facebook/X).
Only sentences with a diacritic density of >=25% were kept to ensure
enough phonetic information for accurate synthesis.
model:
Resemble AI… See the full description on the dataset page: https://huggingface.co/datasets/AymanMansour/New-Lisan-Sudanese-TTS-Dataset.kantipur-news-data
Nepali Speech Dataset (YouTube-sourced)
74 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 74 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-news-data.new-moore-speech-clean
Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French
The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language.
[!NOTE]
⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.new-dioula-speech-clean
Dioula Speech Corpus: A Parallel Audio-Text Dataset for Dioula and French
The Dioula Speech Corpus is a bilingual audio-text corpus designed for research and academic purposes in low-resource speech and language processing. It is intended primarily to support the development of Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models for the Dioula language.
⚠️ Access is gated. To request access, please read the policy below.🛑 TLDR: For safety and traceability reasons… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-dioula-speech-clean.new_taipei_education_department_dataset
海岸阿美語語音資料集
資料集摘要
本資料集收錄台灣海岸阿美語的語音與文字對照,適用於自動語音辨識(ASR)、語音合成(TTS)及原住民族語相關研究。
每筆資料包含一段語音、羅馬拼音轉寫、中文翻譯,以及語別標籤。
項目
說明
語言
海岸阿美語(Amis)
樣本數
33,080
Split
train
音檔格式
MP3
平均時長
約 6.6 秒
時長範圍
1.0 – 117.5 秒
資料欄位
欄位
型別
說明
範例
id
string
唯一識別碼
record1-1@0001.mp3
audio
audio
語音檔
(可於 Dataset Viewer 播放)
duration
float64
音檔長度(秒)
3.288
transcript
string
羅馬拼音轉寫
cecay
translation
string
中文翻譯
一
lang_group
string
語別(中文)
海岸阿美語… See the full description on the dataset page: https://huggingface.co/datasets/allen-1216/new_taipei_education_department_dataset.
