CoolFace
Datasetpublic

VelkroLM/african-languages-catalog

VelkroLM African-language corpus catalog This catalog organizes the current audited African-language publication waves. Published repositories Text corpus: https://huggingface.co/datasets/VelkroLM/african-languages-corpus Personal text mirror: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered Speech wave 1: https://huggingface.co/datasets/VelkroLM/african-languages-speech HPLT source: https://hplt-project.org/datasets/v3.0 WAXAL source:… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-catalog.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes31downloads
Dataset Card

VelkroLM African-language corpus catalog

This catalog organizes the current audited African-language publication waves.

Published repositories

  • Text corpus: https://huggingface.co/datasets/VelkroLM/african-languages-corpus
  • Personal text mirror: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered
  • Speech wave 1: https://huggingface.co/datasets/VelkroLM/african-languages-speech
  • HPLT source: https://hplt-project.org/datasets/v3.0
  • WAXAL source: https://huggingface.co/datasets/google/WaxalNLP

The text corpus contains filtered HPLT v3.0 shards for Hausa, Yoruba, Igbo, Fulfulde, Kanuri, Amharic, Somali, Tigrinya, Wolof, Lingala, Luganda, and Shona. The speech wave contains verified Hausa, Yoruba, and Igbo WAXAL TTS audio/transcript tar shards with audit files.

Always inspect each upstream dataset card and license. This catalog does not grant rights beyond the upstream terms.