BeitTigreAI/tigre-data-wikipedia
Tigre Wikipedia Corpus (tigwiki) Overview This repository houses the Tigre Wikipedia Corpus, a foundational linguistic resource containing all non-template articles from https://tig.wikipedia.org. Tigre is an under-resourced South Semitic language within the Afro-Asiatic family. This dataset serves as a critical component for bridging the digital divide, facilitating the development of Natural Language Processing (NLP) models—including Language Models (LMs)… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-wikipedia.
09
No card is published for this repository, or it could not be fetched from Hugging Face right now.
