datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.waxholmThe Waxholm corpus was collected in 1993 - 1994 at the department of Speech, Hearing and Music (TMH), KTH.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/WaxalNLP.waxal-pseudo
WAXAL Pseudo-Labels (3-model agreement)
PRIVATE working artifact for the Google WAXAL ASR Challenge — not for redistribution.
High-confidence pseudo-labels for the WAXAL unlabeled split, produced by a 3-model agreement cascade:
a clip is kept only when the fine-tuned champion (w2v-BERT-2.0 CTC) and XLS-R-300m agree
(CER ≤ 0.12), and the fine-tuned Omnilingual-ASR-300M independently confirms the champion transcript
(CER ≤ 0.22). omni is architecturally diverse (different… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/waxal-pseudo.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/jessteru/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/youvoi/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/claudefitz/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/novelwolde36/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/BolajiJsAI/WaxalNLP.waxal-linsna
WAXAL Phase-2 Lingala / Shona training corpus (derived)
This corpus was built for the Google WAXAL ASR Challenge on Zindi.
This repository is the exact training corpus behind our Google WAXAL ASR Challenge (Phase 2)
submission: the TSV manifests plus the derived 16 kHz mono FLAC audio that our training configs read.
It exists so that the whole recipe can be rebuilt and audited from one place. The manifests below,
together with the google/WaxalNLP train and validation splits read… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/waxal-linsna.WaxalNLP
Waxal Datasets
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language technology for these underserved languages, and
to serve as a repository for digital preservation.
The Waxal datasets are collections acquired through partnerships with Makerere… See the full description on the dataset page: https://huggingface.co/datasets/Phsntom/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/matrixdose/WaxalNLP.WaxalNLP
Waxal Datasets
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language technology for these underserved languages, and
to serve as a repository for digital preservation.
The Waxal datasets are collections acquired through partnerships with… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/WaxalNLP.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.WaxalNLPr
Waxal NLP Datasets
Overview
This repository hosts a large multilingual speech corpus for 27 African languages, split into two task collections: Automatic Speech Recognition (ASR) — natural speech paired with human transcriptions — and Text-to-Speech (TTS) — single-speaker studio recordings paired with the scripted text that was read aloud. The dataset card and data itself indicate the underlying recordings were gathered through partnerships with Makerere… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/WaxalNLPr.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.waxal-audio-text-retrieval
WAXAL speech–text retrieval (MTEB)
Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from
WAXAL (Google and partners).
Most of these languages have no presence in mteb's existing multilingual audio tasks,
which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and
WaxalT2ARetrieval.
Contents
One config per language, each with id, audio, text, speaker_id, gender,
language. 1,722 utterances total.
code
language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/bdallhrajh371/WaxalNLP.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/ngoloan/WaxalNLP.bambara-tts-waxal
bambara-tts-waxal
Bambara studio speech from the WAXAL corpus — 1,926 recordings, 16 hours, 8 speakers,
44.1 kHz mono.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-tts-waxal", "google_waxal", split="train")
Splits: train, validation, test.
Fields
Field
Description
audio
44.1 kHz mono
text
Transcript
speaker_id
Speaker identifier (8 distinct)
gender
Speaker gender
locale
Locale code
id
Record… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-tts-waxal.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/isaacoluwafemiog/WaxalNLP.serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
This is the first pushed Ghanaian Speech Lab ASR pipeline artifact. It is a
review artifact for the v0.1 Akan ASR pass, not a trained model checkpoint.
Expected future model repo:
teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
What This Artifact Contains
data/manifest.jsonl: harmonized Waxal + GhanaNLP manifest references.
reports/sanitize.json: sanitization report and… See the full description on the dataset page: https://huggingface.co/datasets/teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1.waxal-orm-tts-merged
Waxal Oromo TTS Merged
This dataset merges the human-labeled Oromo ASR split from
google/WaxalNLP with the autolabeled Oromo split from
israel/waxal-autolabled.
For TTS use, the leading [ORM] language tag has been removed from
autolabeled transcriptions. Rows include both text and transcription
with the same cleaned value.
Target repo: b1n1yam/waxal-orm-tts-merged
sna-waxal-unlabeled-tar
WAXAL Shona unlabeled operational TAR dataset
WebDataset packaging of the Shona unlabeled split from
google/WaxalNLP.
config: sna_asr
split: unlabeled
pinned upstream revision: e0a62aaebc61bd5bb8cac17a08d1b42c65551dd2
samples: 85,384
audio: source-encoded bytes preserved without transcoding
The pinned upstream Parquet files remain the recoverable source. This repo is
an operational derivative optimized for sequential streaming. Each example is
a matching audio and .json pair.… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-unlabeled-tar.
