datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-170k-dyula
Dataset Description
Code-170k-dyula is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Dyula, making coding education accessible to Dyula speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Dyula language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-dyula.dyula-tswana_sentence-pairs
Dyula-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Tswana_Sentence-Pairs
Number of Rows: 94328
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-tswana_sentence-pairs.dyula-swati_sentence-pairs
Dyula-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Swati_Sentence-Pairs
Number of Rows: 28528
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-swati_sentence-pairs.dyula-lingala_sentence-pairs
Dyula-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Lingala_Sentence-Pairs
Number of Rows: 57764
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-lingala_sentence-pairs.dyula-sentiments-corpus
Dyula Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Dyula for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 36,382
Positive sentiment: 20884 (57.4%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-sentiments-corpus.dyula-xhosa_sentence-pairs
Dyula-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Xhosa_Sentence-Pairs
Number of Rows: 88194
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-xhosa_sentence-pairs.dyula-somali_sentence-pairs
Dyula-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Somali_Sentence-Pairs
Number of Rows: 112864
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-somali_sentence-pairs.dyula_dataset_3.0dyula-ganda_sentence-pairs
Dyula-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Ganda_Sentence-Pairs
Number of Rows: 77324
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-ganda_sentence-pairs.dyula-english_sentence-pairsdyula-speech-bible
Total duration = ~72 hours
sampling rate = 16KHz
dyula-french_sentence-pairsdyula-rundi_sentence-pairs
Dyula-Rundi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Rundi_Sentence-Pairs
Number of Rows: 72663
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-rundi_sentence-pairs.dyula-kongo_sentence-pairs
Dyula-Kongo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Kongo_Sentence-Pairs
Number of Rows: 42989
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-kongo_sentence-pairs.dyula-english-emotions-corpus
Dyula-english Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Dyula-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-english-emotions-corpus.dyula-yoruba_sentence-pairs
Dyula-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Yoruba_Sentence-Pairs
Number of Rows: 102298
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-yoruba_sentence-pairs.dyula-tsonga_sentence-pairs
Dyula-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Tsonga_Sentence-Pairs
Number of Rows: 82077
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-tsonga_sentence-pairs.dyula-oromo_sentence-pairs
Dyula-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Oromo_Sentence-Pairs
Number of Rows: 45729
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-oromo_sentence-pairs.dyula-kikuyu_sentence-pairs
Dyula-Kikuyu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Kikuyu_Sentence-Pairs
Number of Rows: 34322
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-kikuyu_sentence-pairs.dyula-fon_sentence-pairs
Dyula-Fon_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Fon_Sentence-Pairs
Number of Rows: 41426
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-fon_sentence-pairs.bambara-dyula_sentence-pairs
Bambara-Dyula_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Dyula_Sentence-Pairs
Number of Rows: 23484
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-dyula_sentence-pairs.dyula-kinyarwanda_sentence-pairs
Dyula-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Kinyarwanda_Sentence-Pairs
Number of Rows: 102925… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-kinyarwanda_sentence-pairs.dyula-kimbundu_sentence-pairs
Dyula-Kimbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Kimbundu_Sentence-Pairs
Number of Rows: 25167
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-kimbundu_sentence-pairs.dyula-chichewa_sentence-pairs
Dyula-Chichewa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Chichewa_Sentence-Pairs
Number of Rows: 100921
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-chichewa_sentence-pairs.dinka-dyula_sentence-pairs
Dinka-Dyula_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Dyula_Sentence-Pairs
Number of Rows: 17561
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-dyula_sentence-pairs.dyula-umbundu_sentence-pairs
Dyula-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Umbundu_Sentence-Pairs
Number of Rows: 37912
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-umbundu_sentence-pairs.dyula-tumbuka_sentence-pairs
Dyula-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Tumbuka_Sentence-Pairs
Number of Rows: 69174
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-tumbuka_sentence-pairs.dyula-shona_sentence-pairs
Dyula-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Shona_Sentence-Pairs
Number of Rows: 111253
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-shona_sentence-pairs.dyula-ewe_sentence-pairs
Dyula-Ewe_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Ewe_Sentence-Pairs
Number of Rows: 70901
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-ewe_sentence-pairs.afrikaans-dyula_sentence-pairs
Afrikaans-Dyula_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Dyula_Sentence-Pairs
Number of Rows: 130823
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-dyula_sentence-pairs.
