akhanafer/arabic-to-arabizi
Arabic to Arabizi (Levantine) A small parallel dataset of short, everyday Levantine (Lebanese) Arabic sentences written in Arabic script, each paired with an Arabizi (Latin-script) transliteration. Dataset structure Split Rows train 389 test 44 Total 433 Each row has two fields: arabic: the sentence in Arabic script arabizi: the same sentence in Arabizi Example: arabic arabizi شو عم تعمل؟ sho 3m t3ml? رحت عالبحر مع العيلة. rht… See the full description on the dataset page: https://huggingface.co/datasets/akhanafer/arabic-to-arabizi.
Arabic to Arabizi (Levantine)
A small parallel dataset of short, everyday Levantine (Lebanese) Arabic sentences written in Arabic script, each paired with an Arabizi (Latin-script) transliteration.
Dataset structure
Each row has two fields:
arabic: the sentence in Arabic scriptarabizi: the same sentence in Arabizi
Example:
Transliteration conventions
3= ع2= ء / أ (hamza)sh= ش,kh= خ,gh= غhis used for both ح and ه (7is not used)q= ق- Short vowels are mostly omitted; long vowels are written (
a,y/i,ou/o)
How the data was created
The Arabic sentences and their Arabizi transliterations were generated with OpenAI's GPT-4, then reviewed by the author.
Intended uses
Training or evaluating Arabic ↔ Arabizi transliteration, input methods and keyboards, and dialect-aware NLP.
Limitations
- Small size (433 pairs); short, conversational sentences only.
- Lebanese/Levantine dialect only; not representative of other dialects or Modern Standard Arabic.
- Reflects one Arabizi spelling style. Real-world Arabizi varies a lot between writers (e.g.
7for ح,5for خ, more written vowels).
License
This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You may use, share and adapt it, including commercially, as long as you give appropriate credit.
Citation
If you use this dataset, please credit:
akhanafer, "Arabic to Arabizi (Levantine)", Hugging Face, 2026.
https://huggingface.co/datasets/akhanafer/arabic-to-arabizi