datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-qwen-vl-lingala-qlora-vf
Qwen-VL Lingala OCR Dataset
Description
Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B).
train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ.
test: original, non-augmented images only, held out before any oversampling to avoid data leakage.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf.tajik-lora-qlora-benchmark
Tajik LoRA/QLoRA Benchmark
📊 Description
This benchmark contains the complete results of fine-tuning 15+ language models (from 124M to 7B parameters) on a subset of the Tajik language (1000 sentences from the TajikNLPWorld/tajik-web-corpus).The study compares full fine-tuning versus LoRA/QLoRA, evaluating model quality (perplexity), GPU memory usage, and training time.
Key Findings
GPT‑2 medium (full fine-tuning) achieves the lowest perplexity (3.48), but… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-lora-qlora-benchmark.bashkir-lora-qlora-benchmark
Bashkir LoRA/QLoRA Benchmark
📊 Description
This benchmark contains the complete results of fine-tuning various language models (from 82M to 7B parameters) on the Bashkir language. The study compares the effectiveness of LoRA/QLoRA against full fine-tuning, evaluating model quality (perplexity), GPU memory usage, and training time.
Key Findings
Mistral-7B with QLoRA (r=16) achieved the best performance among 7B models (perplexity 3.79)
LoRA drastically reduces… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lora-qlora-benchmark.
