datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-dialects-gold20
arabic-dialects-gold20
660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized
dialectal orthography, the undiacritized surface form, gold IPA, an engine
draft, an English gloss, machine-verified phonetic feature tags, per-row
verification metadata, and notes citing the dialectological literature that
grounds the row.
Columns (TSV, UTF-8, one file per lect):
id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20.Arabic-Dialects
Arabic Dialects Dataset (Bivalency & Code-Switching)
The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties:
EGY – Egyptian Arabic
GLF – Gulf Arabic
LAV – Levantine Arabic
NOR – North African / Tunisian Arabic
MSA – Modern Standard Arabic
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.arabic-dialects-gold20-code-switch
gold20-code-switch
Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic
lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20).
Each row embeds foreign material in a dialectal Arabic frame: inline
Latin-script English (and French, for the lects whose live contact language
is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi
(Latin-written Arabic with digit gutturals).
Columns (TSV, UTF-8, one file per lect):
id… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20-code-switch.Multi-Arabic-dialectsArabic_Dialects
Dataset Card for Arabic Dialects
Dataset Summary
The Arabic Dialects dataset is a collection of text samples representing multiple spoken Arabic dialects alongside Modern Standard Arabic (MSA). It is designed to help train and evaluate natural language processing (NLP) models on dialect identification, text classification, and understanding regional linguistic variations.
Languages and Dialects Included
Egyptian (EGY)
Gulf (GLF)
Levantine (LEV)… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Arabic_Dialects.organic-gulf-arabic-dialect-dataset
Organic Gulf Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Gulf Arabic (Khaleeji) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-gulf-arabic-dialect-dataset.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.Arabic_dialects_to_MSAorganic-levantine-arabic-dialect-dataset
Organic Levantine Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-levantine-arabic-dialect-dataset.arabic-dialect-text
Arabic Dialectal Text — gathered, lang-coded, IPA-enriched
A deduplicated collection of dialectal Arabic sentences assembled from openly-downloadable
sources, every line tagged with a BCP-47 lang code. Saudi Arabic is the focus, but all
labelled dialects are retained. Built as the text side of a Saudi TTS / phonemizer pipeline.
Files
all.tsv — the corpus: id<TAB>lang<TAB>source<TAB>text.
all.enriched.tsv — adds two phonetic columns:… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialect-text.organic-sudanese-arabic-dialect-dataset
Organic Sudanese Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Sudanese Arabic dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-sudanese-arabic-dialect-dataset.organic-iraqi-arabic-dialect-dataset
Organic Iraqi Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Iraqi Arabic dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately captures… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-iraqi-arabic-dialect-dataset.organic-maghrebi-arabic-dialect-dataset
Organic Maghrebi Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Maghrebi Arabic (Darija) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-maghrebi-arabic-dialect-dataset.arabic_dialectsorganic-egyptian-arabic-dialect-dataset
Organic Egyptian Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-egyptian-arabic-dialect-dataset.arabic-dialect-to-msarafeeq-arabic-medical-dialect
Rafiq Arabic Medical Dialect Dataset
Dataset summary
Rafiq is an Arabic dataset for dialect-to-simple-Arabic normalization/translation of health-related expressions.
The primary task is to convert a colloquial Arabic expression into simpler Arabic while preserving its meaning and avoiding added diagnoses or symptoms.
medical_department is provided as an optional auxiliary text-classification/routing label. It is not the primary task and must not be interpreted as… See the full description on the dataset page: https://huggingface.co/datasets/Mayssm/rafeeq-arabic-medical-dialect.
