umbundu
Datasets
All datasets matching “umbundu”umbundu-emotions-corpus
Umbundu Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Umbundu for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/umbundu-emotions-corpus.umbundu-sentiments-corpus
Umbundu Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Umbundu for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 83,350
Positive sentiment: 48940 (58.7%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/umbundu-sentiments-corpus.akan-umbundu_sentence-pairs
Akan-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Umbundu_Sentence-Pairs
Number of Rows: 22651
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-umbundu_sentence-pairs.swahili-umbundu_sentence-pairs
Swahili-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Umbundu_Sentence-Pairs
Number of Rows: 276674
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-umbundu_sentence-pairs.shona-umbundu_sentence-pairs
Shona-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Umbundu_Sentence-Pairs
Number of Rows: 153467
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-umbundu_sentence-pairs.english-umbundu_sentence-pairs_mt560
English-Umbundu Parallel Dataset
This dataset contains parallel sentences in English and Umbundu (Angola).
Dataset Information
Language Pair: English ↔ Umbundu
Language Code: umb
Country: Angola
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-umbundu_sentence-pairs_mt560.
