dyula
Datasets
All datasets matching “dyula”Code-170k-dyula
Dataset Description
Code-170k-dyula is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Dyula, making coding education accessible to Dyula speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Dyula language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-dyula.dyula-tswana_sentence-pairs
Dyula-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Tswana_Sentence-Pairs
Number of Rows: 94328
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-tswana_sentence-pairs.dyula-swati_sentence-pairs
Dyula-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Swati_Sentence-Pairs
Number of Rows: 28528
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-swati_sentence-pairs.dyula-lingala_sentence-pairs
Dyula-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Lingala_Sentence-Pairs
Number of Rows: 57764
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-lingala_sentence-pairs.dyula-sentiments-corpus
Dyula Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Dyula for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 36,382
Positive sentiment: 20884 (57.4%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-sentiments-corpus.dyula-xhosa_sentence-pairs
Dyula-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Xhosa_Sentence-Pairs
Number of Rows: 88194
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-xhosa_sentence-pairs.
