dyu
Datasets
All datasets matching “dyu”Koumankan_mt_dyu_fr
Koumankan4Dyula: Parallel Dyula - French Dataset for Machine Learning
Overview
The Koumankan4Dyula corpus consists of 10929 pairs of Dioula-French sentences.
This corpus is part of the Koumankan project, which proposes a scalable and cost-effective method for extending the CommonVoice dataset to the Dyula language and other African languages.
Data Splits
Train
73%
8065
Valid
14%
1471
Test
13%
1393
Maintenance
This… See the full description on the dataset page: https://huggingface.co/datasets/uvci/Koumankan_mt_dyu_fr.Code-170k-dyula
Dataset Description
Code-170k-dyula is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Dyula, making coding education accessible to Dyula speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Dyula language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-dyula.dyula-tswana_sentence-pairs
Dyula-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Tswana_Sentence-Pairs
Number of Rows: 94328
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-tswana_sentence-pairs.dyula-lingala_sentence-pairs
Dyula-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Lingala_Sentence-Pairs
Number of Rows: 57764
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-lingala_sentence-pairs.dyute_fireemblem
Dataset of dyute (Fire Emblem)
This is the dataset of dyute (Fire Emblem), containing 160 images and their tags.
The core tags of this character are brown_hair, ponytail, bow, brown_eyes, long_hair, fang, hair_bow, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size
Download
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/dyute_fireemblem.dyula-swati_sentence-pairs
Dyula-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Swati_Sentence-Pairs
Number of Rows: 28528
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-swati_sentence-pairs.
