Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg
Tumbuka Text Corpus - Translated Gutenberg Dataset Description This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania. Dataset Summary Language: Tumbuka (tum) Source: Project Gutenberg Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.
Tumbuka Text Corpus - Translated Gutenberg
Dataset Description
This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania.
Dataset Summary
- Language: Tumbuka (tum)
- Source: Project Gutenberg
- Content: Translated literary classics and various text documents.
- Approximate Size: 1.4 GB
- Approximate Token Count: ~203,606,962 tokens
Repository Structure
The dataset is organized into text files within the translated_gutenberg directory. Each file corresponds to a translated work or a segment of text processed from the original English source.
Intended Use
This corpus is intended for:
- Pre-training and fine-tuning Large Language Models (LLMs).
- Developing machine translation systems for Bantu languages.
- Linguistic research and text analysis in Tumbuka.
Dataset Statistics
- Total Size on Disk: 1,403.74 MB
- Format: Plain text (.txt)
- Encoding: UTF-8
Licensing
The original Gutenberg texts are generally in the public domain in the USA. This translated corpus is released under the Apache License 2.0.
Acknowledgements
Special thanks to the contributors and the automated systems used to generate these translations to improve the digital footprint of the Tumbuka language.
