VelkroLM/african-languages-speech
WAXAL African-language speech — filtered wave 1 This repository contains bounded, filtered WAXAL audio/transcript pairs for Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). Each configuration is split into tar archives containing an audio file plus a JSON record with transcript, language, speaker, source ID, and provenance. The upstream source is google/WaxalNLP and the WAXAL paper is arXiv:2602.02734. Consult the upstream dataset card for the exact component license and… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-speech.
WAXAL African-language speech — filtered wave 1
This repository contains bounded, filtered WAXAL audio/transcript pairs for Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). Each configuration is split into tar archives containing an audio file plus a JSON record with transcript, language, speaker, source ID, and provenance.
The upstream source is google/WaxalNLP and the WAXAL paper is arXiv:2602.02734. Consult the upstream dataset card for the exact component license and attribution requirements; this derivative does not grant rights beyond the upstream terms.
The collector streamed one source row at a time, removed empty/obvious spam text and missing audio, preserved transcript/audio alignment, used a seeded random reservoir sample, and deleted raw staging after successful verification. Each configuration includes audit.json with counts, drop reasons, checksums, and samples.
The remaining WAXAL configurations will be added in paced waves rather than duplicated into a second local copy.
