datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/WaxalNLP.waxal-pseudo
WAXAL Pseudo-Labels (3-model agreement)
PRIVATE working artifact for the Google WAXAL ASR Challenge — not for redistribution.
High-confidence pseudo-labels for the WAXAL unlabeled split, produced by a 3-model agreement cascade:
a clip is kept only when the fine-tuned champion (w2v-BERT-2.0 CTC) and XLS-R-300m agree
(CER ≤ 0.12), and the fine-tuned Omnilingual-ASR-300M independently confirms the champion transcript
(CER ≤ 0.22). omni is architecturally diverse (different… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/waxal-pseudo.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/jessteru/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/youvoi/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/claudefitz/WaxalNLP.waxal-amharic-combinedWaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/novelwolde36/WaxalNLP.waxal-autolabled
Auot-Lableing Waxal unlabeled dataset on Best Multilingual Ethio-ASR models
@article{abdullah2026ethio,
title={Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages},
author={Abdullah, Badr M and Azime, Israel Abebe and Tonja, Atnafu Lambebo and Alabi, Jesujoba O and Alemu, Abel Mulat and Hagos, Eyob G and Balcha, Bontu Fufa and Nerea, Mulubrhan A and Yadeta, Debela Desalegn and Marilign, Dagnachew Mekonnen and others}… See the full description on the dataset page: https://huggingface.co/datasets/israel/waxal-autolabled.waxal-asr-lin_sna_lug-combined-datasetWaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/BolajiJsAI/WaxalNLP.WaxalNLP
Waxal Datasets
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language technology for these underserved languages, and
to serve as a repository for digital preservation.
The Waxal datasets are collections acquired through partnerships with Makerere… See the full description on the dataset page: https://huggingface.co/datasets/Phsntom/WaxalNLP.Wolof-Kallaama-WaxalWaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/matrixdose/WaxalNLP.WaxalNLP
Waxal Datasets
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language technology for these underserved languages, and
to serve as a repository for digital preservation.
The Waxal datasets are collections acquired through partnerships with… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/WaxalNLP.waxal-features-v1
waxal-features-v1
Precomputed Whisper-large-v3 log-mel input_features + tokenized labels
for Google WaxalNLP (Lingala, Shona, Luganda).
Use this to skip FLAC download + feature extraction when fine-tuning
openai/whisper-large-v3 (or any model that consumes the same Whisper-v3 mel /
tokenizer layout).
Contents
Field
Type
Notes
id
string
Clip id
language
string
lin / sna / lug
split
string
Source split tag
input_features
list[list[float16]]… See the full description on the dataset page: https://huggingface.co/datasets/mmwanje/waxal-features-v1.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.WaxalNLPr
Waxal NLP Datasets
Overview
This repository hosts a large multilingual speech corpus for 27 African languages, split into two task collections: Automatic Speech Recognition (ASR) — natural speech paired with human transcriptions — and Text-to-Speech (TTS) — single-speaker studio recordings paired with the scripted text that was read aloud. The dataset card and data itself indicate the underlying recordings were gathered through partnerships with Makerere… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/WaxalNLPr.google_waxal_asr_challenge
WaxalNLP ASR — Cleaned Subset (Lingala, Shona, Luganda)
This dataset is a cleaned, corrected subset of google/WaxalNLP, covering the train and validation splits for three languages:
lin_asr — Lingala
sna_asr — Shona
lug_asr — Luganda
The test split from the original dataset is intentionally excluded.
What was changed
The original transcriptions for these three languages contained a number of errors. A corrected transcription file was applied on top of the… See the full description on the dataset page: https://huggingface.co/datasets/Harcuracy/google_waxal_asr_challenge.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.WaxalNLP
Google WaxalNLP Wolof Re-alignment
Google introduced WAXAL, a new open dataset for 21 African languages, to tackle data scarcity and build inclusive speech technology. However, the Wolof language has experienced alignment issues between the audio files and their transcriptions, making the dataset unusable.
We therefore propose to correct this using a simple and effective approach:
For each audio clip, we generated a transcription using Google Gemini ASR.
For each generated… See the full description on the dataset page: https://huggingface.co/datasets/galsenai/WaxalNLP.waxalnlp-amh-orm-asrwaxal-audio-text-retrieval
WAXAL speech–text retrieval (MTEB)
Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from
WAXAL (Google and partners).
Most of these languages have no presence in mteb's existing multilingual audio tasks,
which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and
WaxalT2ARetrieval.
Contents
One config per language, each with id, audio, text, speaker_id, gender,
language. 1,722 utterances total.
code
language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/bdallhrajh371/WaxalNLP.waxal-lug-clean
Waxal Luganda TTS (Cleaned)
A cleaned version of the Luganda TTS subset from Google's WaxalNLP dataset, preprocessed for fine-tuning text-to-speech models.
What changed from the original?
The original Waxal recordings contain click/pop artifacts at the start and end of audio clips (likely from the recording equipment). These transients degrade TTS model quality during fine-tuning.
This dataset applies Silero VAD (Voice Activity Detection) to precisely detect speech… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/waxal-lug-clean.waxal_nlp_tts
WAXAL NLP TTS
waxal-cleaned-16k
waxal-cleaned-16k
Cleaned google/WaxalNLP ASR configs: heuristics + loose MMS-1b zero-shot CER filter
(<= 0.80), audio re-encoded 16 kHz mono FLAC. Audit trail in _audit/, resume state in
_state/. Generated by waxal_onepass_colab.ipynb.
WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/ngoloan/WaxalNLP.waxal-wolof
