Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf
Qwen-VL Lingala OCR Dataset Description Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B). train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ. test: original, non-augmented images only, held out before any oversampling to avoid data leakage.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf.
Qwen-VL Lingala OCR Dataset
Description
Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B).
train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ.test: original, non-augmented images only, held out before any oversampling to avoid data leakage.
General statistics
Training split
Technical characteristics and content
- Data nature: the corpus consists of scanned textual documents, notably from the SIL Congo collection, converted into images for optical character recognition (OCR) training.
- Multimodality: the corpus is structured in multimodal ChatML format, associating scanned page images with their corresponding text transcriptions.
- Objective: prepare model training for the automatic recognition and transcription of Lingala textual documents.
- Anonymization: the handwritten data included in this process has been anonymized.
License
This dataset is published under the Nwulite Obodo Open Data License (NOODL-1.0), a license designed for the equitable sharing of African datasets.
This dataset, qwenlingaladataset (Congo-digital-service/qwenlingaladataset), was created by Congo Digital Services (CDS SARL) (https://www.congo-digital.com/) in August 2026 in collaboration with the MALOBA community (https://maloba.congo-digital.com/), Radio Rurale, the Service National des Grandes Endémies de Brazzaville, the Faculté des Lettres et des Langues Vivantes, Université Marien Ngouabi (UMNG), the Ministry of Posts, Telecommunications and Digital Economy of the Republic of Congo, and UNDP Congo. It is licensed under the Nwulite Obodo Open Data License (https://licensingafricandatasets.com/nwulite-obodo-license). In exchange for use of this dataset, users from developing countries are required to provide: Use of the dataset (no additional benefit required). Users from high-income countries or commercial entities are additionally required to publicly acknowledge and credit the Maloba Project (UNDP Republic of Congo — language digitalisation initiative) in any publication, product, model, or output derived from this dataset. To fulfil this requirement, contact contact@congo-digital.com or visit https://www.congo-digital.com/contact.
Creators
- Congo Digital Services (CDS SARL) — https://www.congo-digital.com/
- In collaboration with:
- the MALOBA community — https://maloba.congo-digital.com/
- SIL Congo
- the Faculté des Lettres et des Langues Vivantes, Université Marien Ngouabi (UMNG)
- the Ministry of Posts, Telecommunications and Digital Economy of the Republic of Congo
- UNDP Congo
- Created in August 2026
