Rundi
Datasets
All datasets matching “Rundi”rundi-sentiments-corpus
Rundi Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Rundi for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 372,663
Positive sentiment: 209740 (56.3%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-sentiments-corpus.english-rundi_sentence-pairs
English-Rundi_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Rundi_Sentence-Pairs
File Size: 94265941 bytes
Languages: English, English
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-rundi_sentence-pairs.Code-170k-rundi
Dataset Description
Code-170k-rundi is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Rundi, making coding education accessible to Rundi speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Rundi language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-rundi.rundi-tumbuka_sentence-pairs
Rundi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Tumbuka_Sentence-Pairs
Number of Rows: 194527
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-tumbuka_sentence-pairs.ewe-rundi_sentence-pairs
Ewe-Rundi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Rundi_Sentence-Pairs
Number of Rows: 197511
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-rundi_sentence-pairs.english-rundi_sentence-pairs_mt560
English-Rundi Parallel Dataset
This dataset contains parallel sentences in English and Rundi (Burundi).
Dataset Information
Language Pair: English ↔ Rundi
Language Code: run
Country: Burundi
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-rundi_sentence-pairs_mt560.
