CoolFace
Datasetpublic

Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf

Qwen-VL Lingala OCR Dataset Description Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B). train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ. test: original, non-augmented images only, held out before any oversampling to avoid data leakage.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf.

sourceHugging Faceotherupdated 20d agoView on Hugging Face
0likes92downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf · CoolFace