datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs_xho
FLEURS -- isiXhosa (xh_za)
Re-mirrored from google/fleurs, config
xh_za. n-way parallel read speech built on FLORES-101 text -- an
evaluation-sized corpus (~19 h), not training scale, but the
de facto African-language ASR/TTS benchmark (used in Whisper, MMS, SeamlessM4T,
USM papers).
Licence
CC BY 4.0 -- inherited unchanged from the source.
What changed from the source
audio peak-normalized per clip (see below); sample rate and encoding otherwise… See the full description on the dataset page: https://huggingface.co/datasets/simpra/fleurs_xho.simple-escwa
🗣️ Simple-ESCWA: A simpler version of ESCWA-CS Corpus
The ESCWA-CS Corpus was collected over two days of meetings of the United Nations Economic and Social Commission for Western Asia (ESCWA) held in 2019.It contains intra-sentential code-switching between Arabic and English, with some speakers—particularly from Algeria, Tunisia, and Morocco—alternating between Arabic and French.
The dataset spans approximately 2.8 hours of speech, featuring dialectal Arabic and a Code Mixing Index… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/simple-escwa.
