VelkroLM/african-languages-catalog
VelkroLM African-language corpus catalog This catalog organizes the current audited African-language publication waves. Published repositories Text corpus: https://huggingface.co/datasets/VelkroLM/african-languages-corpus Personal text mirror: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered Speech wave 1: https://huggingface.co/datasets/VelkroLM/african-languages-speech HPLT source: https://hplt-project.org/datasets/v3.0 WAXAL source:… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-catalog.
VelkroLM African-language corpus catalog
This catalog organizes the current audited African-language publication waves.
Published repositories
- Text corpus: https://huggingface.co/datasets/VelkroLM/african-languages-corpus
- Personal text mirror: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered
- Speech wave 1: https://huggingface.co/datasets/VelkroLM/african-languages-speech
- HPLT source: https://hplt-project.org/datasets/v3.0
- WAXAL source: https://huggingface.co/datasets/google/WaxalNLP
The text corpus contains filtered HPLT v3.0 shards for Hausa, Yoruba, Igbo, Fulfulde, Kanuri, Amharic, Somali, Tigrinya, Wolof, Lingala, Luganda, and Shona. The speech wave contains verified Hausa, Yoruba, and Igbo WAXAL TTS audio/transcript tar shards with audit files.
Always inspect each upstream dataset card and license. This catalog does not grant rights beyond the upstream terms.
