datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sango-vocabulary
Sango Vocabulary Dataset
Dataset Description
An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources.
This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.Code-170k-sango
Dataset Description
Code-170k-sango is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Sango, making coding education accessible to Sango speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Sango language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/Code-170k-sango.Code-170k-sango
Dataset Description
Code-170k-sango is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Sango, making coding education accessible to Sango speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Sango language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-sango.sango-french-bible-parallel
SFPC: Sango-French Parallel Corpus
The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language.
Associated resources:
Model: alaminerca/nllb-sango-french
Demo: Sango-French Translator
Paper: SangoNMT: Parameter-Efficient Domain Adaptation of… See the full description on the dataset page: https://huggingface.co/datasets/alaminerca/sango-french-bible-parallel.sango-emotions-corpus
Sango Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Sango for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples: 119… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/sango-emotions-corpus.english-sango_sentence-pairs_mt560
English-Sango Parallel Dataset
This dataset contains parallel sentences in English and Sango (Central African Republic).
Dataset Information
Language Pair: English ↔ Sango
Language Code: sag
Country: Central African Republic
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-sango_sentence-pairs_mt560.sango-sentiments-corpus
Sango Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Sango for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 119,538
Positive sentiment: 73005 (61.1%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/sango-sentiments-corpus.english-sango_sentence-pairs
English-Sango_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Sango_Sentence-Pairs
File Size: 348342520 bytes
Languages: English, English
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-sango_sentence-pairs.french-sango_sentence-pairs
