sango
Datasets
All datasets matching “sango”sango-vocabulary
Sango Vocabulary Dataset
Dataset Description
An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources.
This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.Code-170k-sango
Dataset Description
Code-170k-sango is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Sango, making coding education accessible to Sango speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Sango language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/Code-170k-sango.Code-170k-sango
Dataset Description
Code-170k-sango is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Sango, making coding education accessible to Sango speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Sango language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-sango.sangonomiya_kokomi_genshin
Dataset of sangonomiya_kokomi/珊瑚宮心海/珊瑚宫心海 (Genshin Impact)
This is the dataset of sangonomiya_kokomi/珊瑚宮心海/珊瑚宫心海 (Genshin Impact), containing 500 images and their tags.
The core tags of this character are long_hair, pink_hair, multicolored_hair, bow-shaped_hair, purple_eyes, bow, gradient_hair, blunt_bangs, very_long_hair, blue_hair, breasts, hair_ornament, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/sangonomiya_kokomi_genshin.sango-french-bible-parallel
SFPC: Sango-French Parallel Corpus
The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language.
Associated resources:
Model: alaminerca/nllb-sango-french
Demo: Sango-French Translator
Paper: SangoNMT: Parameter-Efficient Domain Adaptation of… See the full description on the dataset page: https://huggingface.co/datasets/alaminerca/sango-french-bible-parallel.sango-emotions-corpus
Sango Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Sango for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples: 119… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/sango-emotions-corpus.
