arabizi
Datasets
All datasets matching “arabizi”nilechat-arabizi-egy
Arabizi-Egypt: A Resource for Advancing Egyptian Arabic Language Models
Arabizi-Egypt is a substantial dataset specifically developed to foster the creation and improvement of language models for the Egyptian Arabic dialect. This resource consists of a converted Egyptian Arabic dialect in Arabic script to Arabizi from a sample taken from both https://huggingface.co/datasets/UBC-NLP/fineweb-edu-Egypt and https://huggingface.co/datasets/UBC-NLP/LHV-Egypt.
Dataset Snapshot:… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/nilechat-arabizi-egy.arabizi-kit-corpus
ArabiziKit Corpus
LLM-annotated sentences of real social-media Arabizi (Arabic written in
Latin letters and digits) with Arabic-script references and dialect tags,
stratified into train / dev / test splits.
355 annotated sentences harvested from public Hugging Face datasets
(filtered for Arabizi, URLs/mentions/hashtags/emoji stripped, split into
sentences)
323 usable rows after dropping empty references: train 226 / dev 48 /
test 49, stratified by dialect
Inter-annotator… See the full description on the dataset page: https://huggingface.co/datasets/rabeeeehh/arabizi-kit-corpus.arabizi_toxic_harassment_words
license: other
task_categories:
- text-classification
task_ids:
- hate-speech-detection
language:
- ar
tags:
- arabizi
- dialectal-arabic
- toxicity
- harassment
- social-media-safety
- red-teaming
size_categories:
- n<1K
🛡️ Pro-Grade Arabizi 250- Sentences - Safety & Moderation Dataset
Purchase Full Access
👉 Buy the Complete Dataset on Gumroad
Arabizi Toxicity & Content Moderation Dataset
A dataset containing Arabizi (3araby) toxic, abusive, and… See the full description on the dataset page: https://huggingface.co/datasets/Arabicc/arabizi_toxic_harassment_words.Arabizi_TransliterationDarija-Text-Ar-Arabiziarabic-to-arabizi
Arabic to Arabizi (Levantine)
A small parallel dataset of short, everyday Levantine (Lebanese) Arabic sentences written in Arabic script, each paired with an Arabizi (Latin-script) transliteration.
Dataset structure
Split
Rows
train
389
test
44
Total
433
Each row has two fields:
arabic: the sentence in Arabic script
arabizi: the same sentence in Arabizi
Example:
arabic
arabizi
شو عم تعمل؟
sho 3m t3ml?
رحت عالبحر مع العيلة.
rht… See the full description on the dataset page: https://huggingface.co/datasets/akhanafer/arabic-to-arabizi.
