CoolFace
Datasetpublic

galsenai/WaxalNLP

Google WaxalNLP Wolof Re-alignment Google introduced WAXAL, a new open dataset for 21 African languages, to tackle data scarcity and build inclusive speech technology. However, the Wolof language has experienced alignment issues between the audio files and their transcriptions, making the dataset unusable. We therefore propose to correct this using a simple and effective approach: For each audio clip, we generated a transcription using Google Gemini ASR. For each generated… See the full description on the dataset page: https://huggingface.co/datasets/galsenai/WaxalNLP.

sourceHugging Faceupdated 6mo agoView on Hugging Face
10likes183downloads
17 commits on main
8af9cfc6mo ago

Update README.md

derguene
8104ca06mo ago

Update README.md

derguene
0a208db6mo ago

Update README.md

derguene
f72fa306mo ago

Added the speakers genders

derguene
3ae58d16mo ago

Update README.md

derguene
94573c06mo ago

Update README.md

derguene
49334f86mo ago

Update README.md

derguene
2631c446mo ago

Update README.md

derguene
37192e76mo ago

Update README.md

derguene
92b79c86mo ago

Update README.md

derguene
61eed2f6mo ago

Update README.md

derguene
c8f78826mo ago

Update README.md

derguene
0dacdc76mo ago

Update README.md

derguene
23c277c6mo ago

Update README.md

derguene
c89469e6mo ago

Update README.md

derguene
71801d06mo ago

Initial upload of the corrected and cleaned version of the Google/WaxalNLP dataset

derguene
d7c75476mo ago

initial commit

derguene