datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.Lingala_100hrs
Lingala 100hrs
110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated
from three publicly available CC-BY-4.0 corpora for ASR research.
Composition
Counts from a full-pass audit on 2026-07-09:
Source
Upstream location
Rows
Splits
AfriVoice (Lingala)
https://huggingface.co/datasets/DigitalUmuganda/AfriVoice
17,544
train (16,144), validation (915), test (485)
LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.tts_lingala_malelingala-speech-datasetqwen-vl-lingala-dataset-augmented
Augmentation
This dataset derives from dataset-qwen-vl-lingala-qlora-vf (417 train / 50 test) through an augmentation step applied to the training image-text pairs, bringing the training volume to 884 examples.
Augmentation method: the 467 additional training examples compared to the source (417 → 884) are obtained mainly through controlled degradation of the input image — noise, brightness/contrast variation, light blur — rather than through synthetic content generation or text… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/qwen-vl-lingala-dataset-augmented.dataset-qwen-vl-lingala-qlora-vf
Qwen-VL Lingala OCR Dataset
Description
Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B).
train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ.
test: original, non-augmented images only, held out before any oversampling to avoid data leakage.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf.Lingala-Alpaca-Dataset-Augmented
Lingala Alpaca LLM Dataset (Augmented)
This dataset is an augmented version of the original dataset available at Congo-digital-service/Dataset-Lingala-Alpaca-LLM, with a larger and more diverse training set than the source, while preserving the original dataset's purpose and structure.
Dataset Structure
Same fields as the source dataset: instruction, input, output, source_project, split into train and test. See the source dataset card for the category/domain… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/Lingala-Alpaca-Dataset-Augmented.Dataset-Lingala-Alpaca-LLM
Lingala Alpaca LLM Dataset
This dataset contains human-generated instruction-following data in Lingala, inspired by the Alpaca dataset format, covering five task categories (formal, pedagogical, summary, urban, translation) across five public-interest domains, with a balanced representation of stylistic categories and instruction types.
General Statistics
Total volume: 6,664 lines in total.
Number of unique source texts: 1,772.
Structure: The dataset is divided… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/Dataset-Lingala-Alpaca-LLM.wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3lingala_20hralpaca-lingala-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-lingala-cleaned.lingala-sentiments-corpus
Lingala Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Lingala for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 427,979
Positive sentiment: 251923 (58.9%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-sentiments-corpus.Code-170k-lingala
Dataset Description
Code-170k-lingala is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Lingala, making coding education accessible to Lingala speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Lingala language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-lingala.lingala-emotions-corpus
Lingala Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Lingala for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-emotions-corpus.english-lingala_sentence-pairs_mt560
English-Lingala Parallel Dataset
This dataset contains parallel sentences in English and Lingala (Democratic Republic of the Congo).
Dataset Information
Language Pair: English ↔ Lingala
Language Code: lin
Country: Democratic Republic of the Congo
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-lingala_sentence-pairs_mt560.french-lingala_sentence-pairs
French-Lingala_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: French-Lingala_Sentence-Pairs
File Size: 78580422 bytes
Languages: French, French
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/french-lingala_sentence-pairs.kamba-lingala_sentence-pairs
Kamba-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Lingala_Sentence-Pairs
Number of Rows: 50317
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-lingala_sentence-pairs.lingala_10hrLingala_bmd_testlingala-swahili_sentence-pairs
Lingala-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Swahili_Sentence-Pairs
Number of Rows: 391907
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-swahili_sentence-pairs.lingala_5hrbemba-lingala_sentence-pairs
Bemba-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Lingala_Sentence-Pairs
Number of Rows: 140378
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-lingala_sentence-pairs.lingala_real_eval_benchmark_croped
Lingala Real Eval Benchmark Cropped
Small cropped real-world Lingala audio benchmark for testing BLI ASR 0.
The dataset contains short audio clips cropped from longer real-world files, covering different domains such as news, catechesis, comedy, cartoon and interview speech.
This dataset is intended for quick qualitative ASR testing and human review. It is not a training dataset.
english-lingala_sentence-pairs
English-Lingala_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Lingala_Sentence-Pairs
File Size: 307050723 bytes
Languages: English, English
Dataset Description
The dataset contains sentence pairs… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-lingala_sentence-pairs.lingala-asr-dataset
Lingala Automatic Speech Recognition ( Text To Speech & Speech To Text Dataset )
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
kongo-lingala-swahili-french-english-pairs-bibledyula-lingala_sentence-pairs
Dyula-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Lingala_Sentence-Pairs
Number of Rows: 57764
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-lingala_sentence-pairs.alpaca_lingala_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_lingala_taco.Code-170k-lingala
Dataset Description
Code-170k-lingala is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Lingala, making coding education accessible to Lingala speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Lingala language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/tokossapp/Code-170k-lingala.
