datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
albanian-error-augmentation
Albanian Controlled Error Augmentation Dataset
Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation.
Each row is an (incorrect → correct) pair with an explicit error_type label.
Error types
error_type
Description
missing_diacritic
Missing ë / ç
c_q_confusion
Confusion between ç / q / c
digraph_reduction
Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.albanian-lang-gemma-format
Included Datasets
This dataset has unique data, but also includes all the data from the below:
https://huggingface.co/datasets/olsi8/open-perfectblend-alb
