datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
audio-mia-batch-20260312
Audio MIA Batch 20260312
This dataset contains 6,998 audio files (64 GB) downloaded from YouTube videos.
Dataset Structure
Each row contains:
audio: Audio bytes (playable in the dataset viewer)
video_id: YouTube video ID
category: Content category
source_term: Search term used
query: Full search query
title: Video title
url: YouTube URL
uploader: Channel name
channel_id: YouTube channel ID
upload_date: Upload date (YYYY-MM-DD)
duration: Video duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/audio-mia-batch-20260312.kiswahili-tts-dataset
Kiswahili Tts Dataset
Dataset Description
Kiswahili (Swahili) TTS dataset combining two sub-collections: (1) the 'A Kiswahili Dataset for Development of Text-To-Speech System' corpus (Kiswa-XXXXX studio recordings with Biblical and general-domain text), and (2) crowdsourced Swahili speech (swh_XXX recordings at 16kHz, studio quality). Covers topics relevant to East Africa.
Languages
Language: Kiswahili / Swahili (sw)
BCP-47: sw
Source tag
kiswahili… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/kiswahili-tts-dataset.buaiir_voice_jap
BUAIIR Japadhola Voice (Bateesa/buaiir_voice_jap)
Separate dataset — structured student read-speech in Japadhola (Adhola, ISO 639-3: adh)
from Busitema University cohorts (Phase-2 batches v, e, v2, e2).
This repo does not include Papoli community recordings; those are published separately at
Bateesa/popolivoice.
Summary
Property
Value
Recordings
10,332 utterances
Duration
~27.6 hours
Language
Japadhola (adh)
Collection
Structured read speech… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/buaiir_voice_jap.tobydata-tts-dataset
Tobydata Tts Dataset
Dataset Description
Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on tailoring, fashion, and vocational training topics. Recorded via mobile application.
Languages
Language: Luganda (lg)
BCP-47: lg
Source tag
tobydata — value of the source column in every row.
Dataset Structure
Column
Type
Description
audio
Audio
Raw WAV audio at original recording… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/tobydata-tts-dataset.rw-tts-dataset
Rw Tts Dataset
Dataset Description
Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, collected in Rwanda.
Languages
Language: Kinyarwanda (rw)
BCP-47: rw
Source tag
rw — identifies the origin of each sample in the source column.
Dataset Structure
Column
Type
Description
audio
Audio
Raw WAV audio at original recording frequency
text
string
Transcription of the spoken content… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/rw-tts-dataset.preacher-tts-dataset
Preacher TTS Dataset
Dataset Description
English speech dataset containing preacher/sermon audio recordings (grace_N.wav)
segmented and transcribed. Audio recorded at 24 kHz stereo.
Languages
Language: English
BCP-47: en
Source tag
preacher — value of the source column in every row.
Dataset Structure
Column
Type
Description
audio
Audio
Raw WAV audio at original recording frequency
text
string
Transcription of the spoken… See the full description on the dataset page: https://huggingface.co/datasets/Bateesa/preacher-tts-dataset.fluers-mn
fleurs-mn
Mongolian speech recognition dataset , recombined and split into a 90% train and 10% test set.
Dataset Statistics
Total samples: 4,428Total duration: 15h 33m 10s (15.55 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
3,985
13h 56m 57s (13.95 h)
12.60 s
test
443
1h 36m 13s (1.60 h)
13.03 s
omni_source
Omni source dataset
Mongolian audio/text pairs used as source material for the omni training set.
Dataset Statistics
Total samples: 69
Total duration: 0h 15m 47s (0.26 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
69
0h 15m 47s (0.26 h)
13.73 s
Single split (no held-out test set) -- this is source/reference data, not a
benchmark split.
