CoolFace
Datasetpublic

Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg

Tumbuka Text Corpus - Translated Gutenberg Dataset Description This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania. Dataset Summary Language: Tumbuka (tum) Source: Project Gutenberg Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes757downloads
Dataset Card

Tumbuka Text Corpus - Translated Gutenberg

Dataset Description

This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania.

Dataset Summary

  • Language: Tumbuka (tum)
  • Source: Project Gutenberg
  • Content: Translated literary classics and various text documents.
  • Approximate Size: 1.4 GB
  • Approximate Token Count: ~203,606,962 tokens

Repository Structure

The dataset is organized into text files within the translated_gutenberg directory. Each file corresponds to a translated work or a segment of text processed from the original English source.

Intended Use

This corpus is intended for:

  • Pre-training and fine-tuning Large Language Models (LLMs).
  • Developing machine translation systems for Bantu languages.
  • Linguistic research and text analysis in Tumbuka.

Dataset Statistics

  • Total Size on Disk: 1,403.74 MB
  • Format: Plain text (.txt)
  • Encoding: UTF-8

Licensing

The original Gutenberg texts are generally in the public domain in the USA. This translated corpus is released under the Apache License 2.0.

Acknowledgements

Special thanks to the contributors and the automated systems used to generate these translations to improve the digital footprint of the Tumbuka language.