albanian
Datasets
All datasets matching “albanian”albanian-common-voice-v1.0albanian-error-augmentation
Albanian Controlled Error Augmentation Dataset
Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation.
Each row is an (incorrect → correct) pair with an explicit error_type label.
Error types
error_type
Description
missing_diacritic
Missing ë / ç
c_q_confusion
Confusion between ç / q / c
digraph_reduction
Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.albanian-medical-exams-chemistry-mcq-270
Albanian Medical Exams MCQ Dataset
This dataset contains 270 multiple-choice questions from Albanian medical exams.
Usage
This dataset contains 270 multiple-choice questions from Albanian medical exams, specifically under
Fondet e pyetjeve >> Profili Kimi, sections available here: https://qsha.gov.al/porvimi-i-informatizuar-i-mjekesise/.
License
The questions in this dataset have been extracted from the official digital medical exams provided by the Ministry… See the full description on the dataset page: https://huggingface.co/datasets/marjpri/albanian-medical-exams-chemistry-mcq-270.AlbanianSpeechAlbanian_WikiOrca
Dataset Card for "Albanian_WikiOrca"
More Information needed
albanian-synthetic
Albanian Synthetic Q&A Dataset
NOTE: DATASET MAY NOT BE WITH ACCURATE INFORMATION, AS IT IS AI-GENERATED.
A high-quality synthetic dataset of Albanian question-answer pairs covering diverse topics, generated using Mistral-Medium via the Le Platforme API.
⚠️ Note: Code for synthetic data generation can be found in the project repository.
Dataset Details
Language: Albanian (Shqip)
Format: CSV (UTF-8 encoded)
Columns:
Prompt: Question in Albanian
Response: Detailed… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/albanian-synthetic.
