CoolFace
Datasetpublic

Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf

Qwen-VL Lingala OCR Dataset Description Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B). train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ. test: original, non-augmented images only, held out before any oversampling to avoid data leakage.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf.

sourceHugging Faceotherupdated 20d agoView on Hugging Face
0likes92downloads
Dataset Card

Qwen-VL Lingala OCR Dataset

Description

Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B).

  • —train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ.
  • —test: original, non-augmented images only, held out before any oversampling to avoid data leakage.

General statistics

MetricValue
Total volume1,562 training/test line pairs
FormatJSONL export (multimodal ChatML)
Data typeScanned pages (anonymized recto files)

Training split

SplitLines
Train1,480
Test82

Technical characteristics and content

  • —Data nature: the corpus consists of scanned textual documents, notably from the SIL Congo collection, converted into images for optical character recognition (OCR) training.
  • —Multimodality: the corpus is structured in multimodal ChatML format, associating scanned page images with their corresponding text transcriptions.
  • —Objective: prepare model training for the automatic recognition and transcription of Lingala textual documents.
  • —Anonymization: the handwritten data included in this process has been anonymized.

License

This dataset is published under the Nwulite Obodo Open Data License (NOODL-1.0), a license designed for the equitable sharing of African datasets.

This dataset, qwenlingaladataset (Congo-digital-service/qwenlingaladataset), was created by Congo Digital Services (CDS SARL) (https://www.congo-digital.com/) in August 2026 in collaboration with the MALOBA community (https://maloba.congo-digital.com/), Radio Rurale, the Service National des Grandes Endémies de Brazzaville, the Faculté des Lettres et des Langues Vivantes, Université Marien Ngouabi (UMNG), the Ministry of Posts, Telecommunications and Digital Economy of the Republic of Congo, and UNDP Congo. It is licensed under the Nwulite Obodo Open Data License (https://licensingafricandatasets.com/nwulite-obodo-license). In exchange for use of this dataset, users from developing countries are required to provide: Use of the dataset (no additional benefit required). Users from high-income countries or commercial entities are additionally required to publicly acknowledge and credit the Maloba Project (UNDP Republic of Congo — language digitalisation initiative) in any publication, product, model, or output derived from this dataset. To fulfil this requirement, contact contact@congo-digital.com or visit https://www.congo-digital.com/contact.

Creators

  • —Congo Digital Services (CDS SARL) — https://www.congo-digital.com/
  • —In collaboration with:
  • —the MALOBA community — https://maloba.congo-digital.com/
  • —SIL Congo
  • —the Faculté des Lettres et des Langues Vivantes, Université Marien Ngouabi (UMNG)
  • —the Ministry of Posts, Telecommunications and Digital Economy of the Republic of Congo
  • —UNDP Congo
  • —Created in August 2026