datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tumbuka_Text-SpeechTumbuka_Text_Corpus_Translated_Gutenberg
Tumbuka Text Corpus - Translated Gutenberg
Dataset Description
This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania.
Dataset Summary
Language: Tumbuka (tum)
Source: Project Gutenberg
Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.Tumbuka_language
About this dataset
This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia.
Usecases
mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania).
Formats
The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_language.Code-170k-tumbuka
Dataset Description
Code-170k-tumbuka is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Tumbuka, making coding education accessible to Tumbuka speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Tumbuka language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-tumbuka.Tumbuka_Instruction_Tuningtumbuka-sentiments-corpus
Tumbuka Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Tumbuka for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 190,542
Positive sentiment: 109522 (57.5%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-sentiments-corpus.rundi-tumbuka_sentence-pairs
Rundi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Tumbuka_Sentence-Pairs
Number of Rows: 194527
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-tumbuka_sentence-pairs.Tumbuka_Continuous_Next-Token_Prediction_Datasettigrinya-tumbuka_sentence-pairs
Tigrinya-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Tumbuka_Sentence-Pairs
Number of Rows: 152916… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-tumbuka_sentence-pairs.english-tumbuka_sentence-pairs_mt560
English-Tumbuka Parallel Dataset
This dataset contains parallel sentences in English and Tumbuka (Malawi).
Dataset Information
Language Pair: English ↔ Tumbuka
Language Code: tum
Country: Malawi
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-tumbuka_sentence-pairs_mt560.igbo-tumbuka_sentence-pairs
Igbo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tumbuka_Sentence-Pairs
Number of Rows: 133589
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tumbuka_sentence-pairs.malawi-chichewa-tumbuka-corpusfon-tumbuka_sentence-pairs
Fon-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Tumbuka_Sentence-Pairs
Number of Rows: 73794
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-tumbuka_sentence-pairs.swahili-tumbuka_sentence-pairs
Swahili-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Tumbuka_Sentence-Pairs
Number of Rows: 499551
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-tumbuka_sentence-pairs.english-tumbuka_sentence-pairs
English-Tumbuka_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Tumbuka_Sentence-Pairs
File Size: 80213128 bytes
Languages: English, English
Dataset Description
The dataset contains sentence pairs… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-tumbuka_sentence-pairs.Tumbuka_Translated_TinyStoriessomali-tumbuka_sentence-pairs
Somali-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tumbuka_Sentence-Pairs
Number of Rows: 179589
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tumbuka_sentence-pairs.kamba-tumbuka_sentence-pairs
Kamba-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Tumbuka_Sentence-Pairs
Number of Rows: 63077
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-tumbuka_sentence-pairs.tumbuka-emotions-corpus
Tumbuka Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Tumbuka for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-emotions-corpus.dinka-tumbuka_sentence-pairs
Dinka-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Tumbuka_Sentence-Pairs
Number of Rows: 23524
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-tumbuka_sentence-pairs.tumbuka-twi_sentence-pairs
Tumbuka-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Twi_Sentence-Pairs
Number of Rows: 176324
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-twi_sentence-pairs.pedi-tumbuka_sentence-pairs
Pedi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Tumbuka_Sentence-Pairs
Number of Rows: 101945
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-tumbuka_sentence-pairs.nuer-tumbuka_sentence-pairs
Nuer-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Tumbuka_Sentence-Pairs
Number of Rows: 19926
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-tumbuka_sentence-pairs.kikuyu-tumbuka_sentence-pairs
Kikuyu-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Tumbuka_Sentence-Pairs
Number of Rows: 54854
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-tumbuka_sentence-pairs.tumbuka-umbundu_sentence-pairs
Tumbuka-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Umbundu_Sentence-Pairs
Number of Rows: 99125
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-umbundu_sentence-pairs.oromo-tumbuka_sentence-pairs
Oromo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Tumbuka_Sentence-Pairs
Number of Rows: 73849
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-tumbuka_sentence-pairs.kongo-tumbuka_sentence-pairs
Kongo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Tumbuka_Sentence-Pairs
Number of Rows: 93758
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-tumbuka_sentence-pairs.hausa-tumbuka_sentence-pairs
Hausa-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Tumbuka_Sentence-Pairs
Number of Rows: 260765
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-tumbuka_sentence-pairs.tsonga-tumbuka_sentence-pairs
Tsonga-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Tumbuka_Sentence-Pairs
Number of Rows: 203106
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-tumbuka_sentence-pairs.tswana-tumbuka_sentence-pairs
Tswana-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tswana-Tumbuka_Sentence-Pairs
Number of Rows: 187262
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tswana-tumbuka_sentence-pairs.
