datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
darja-blindspot-eval
Blind Spot Evaluation: Algerian Darja, Arabizi, and French Code-Switching
Author: Khadija Abderrahmane
Model evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.5B parameters)
1. The blind spot
Nearly every widely used Arabic NLP benchmark — ArabicMMLU, ARLUE, AraSentiment, and most Arabic instruction-tuning datasets — is built almost entirely on Modern Standard Arabic (MSA), with limited coverage of major spoken dialects (Egyptian, Gulf, Levantine). Algerian Darja, the… See the full description on the dataset page: https://huggingface.co/datasets/khadidjaabderrahmane/darja-blindspot-eval.darja-tounsi1darjaAI_datasetmsa-darja-pairs-completemsa-darja-pairs
