Mwanzau/Tumbuka_language
About this dataset This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia. Usecases mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania). Formats The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_language.
About this dataset
This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia.
Usecases
mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania).
Formats
The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt, .json and .csv, you just need to pick the one that suites your needs.
Contents
The datasets, are also covering a wide range of topics, from:
- life style/cultural greetings,
- history,
- agriculture,
- biology,
- geography,
- astronomy
- Note: So many more will be added in the future to atleast cover all major topics in life. My main goal is to make it easier for the AI models to understand the Tumbuka language from a variety of topics, most of which most Bantu languages lack the vocabulary of.
The Main Tumbuka Versions that are currently in this Dataset Repo are
- Malawian Dialect
- Zambian Dialect
>> I hope this dataset helps you achieve your goal.! 🙂️
