CoolFace
Datasetpublic

Mwanzau/Tumbuka_language

About this dataset This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia. Usecases mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania). Formats The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_language.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes108downloads
Dataset Card

About this dataset

This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia.

Usecases

mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania).

Formats

The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt, .json and .csv, you just need to pick the one that suites your needs.

Contents

The datasets, are also covering a wide range of topics, from:

  1. 1.life style/cultural greetings,
  2. 2.history,
  3. 3.agriculture,
  4. 4.biology,
  5. 5.geography,
  6. 6.astronomy
  • Note: So many more will be added in the future to atleast cover all major topics in life. My main goal is to make it easier for the AI models to understand the Tumbuka language from a variety of topics, most of which most Bantu languages lack the vocabulary of.

The Main Tumbuka Versions that are currently in this Dataset Repo are

  1. 1.Malawian Dialect
  2. 2.Zambian Dialect

>> I hope this dataset helps you achieve your goal.! 🙂️