Congo-digital-service/qwen-vl-lingala-dataset-augmented
Augmentation This dataset derives from dataset-qwen-vl-lingala-qlora-vf (417 train / 50 test) through an augmentation step applied to the training image-text pairs, bringing the training volume to 884 examples. Augmentation method: the 467 additional training examples compared to the source (417 → 884) are obtained mainly through controlled degradation of the input image — noise, brightness/contrast variation, light blur — rather than through synthetic content generation or text… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/qwen-vl-lingala-dataset-augmented.
Augmentation
This dataset derives from `dataset-qwen-vl-lingala-qlora-vf` (417 train / 50 test) through an augmentation step applied to the training image-text pairs, bringing the training volume to 884 examples.
Augmentation method: the 467 additional training examples compared to the source (417 → 884) are obtained mainly through controlled degradation of the input image — noise, brightness/contrast variation, light blur — rather than through synthetic content generation or text rephrasing: the transcription associated with each augmented image remains strictly identical to that of the original image it derives from. A targeted oversampling is additionally applied to lines containing the two lingala-specific characters ɔ (U+0254) and ɛ (U+025B), underrepresented in the source corpus, to improve their recognition by the model.
The test split (50 examples) remains identical to the source's — not augmented — so as not to bias evaluation.
Purpose of the technique: to bring training conditions closer to the real-world variability of scanned documents (uneven scan quality, variable lighting, sensor noise) without ever altering the textual ground truth.
License and attribution: identical to those of the source dataset — augmentation does not create new rights, it derives from the original corpus.
