datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-dialects-gold20
arabic-dialects-gold20
660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized
dialectal orthography, the undiacritized surface form, gold IPA, an engine
draft, an English gloss, machine-verified phonetic feature tags, per-row
verification metadata, and notes citing the dialectological literature that
grounds the row.
Columns (TSV, UTF-8, one file per lect):
id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20.Arabic-Dialects
Arabic Dialects Dataset (Bivalency & Code-Switching)
The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties:
EGY – Egyptian Arabic
GLF – Gulf Arabic
LAV – Levantine Arabic
NOR – North African / Tunisian Arabic
MSA – Modern Standard Arabic
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.shironaam
Dataset Card for Shironaam Corpus
Dataset Summary
Automatic headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers.
However, the performance of headline generation systems remains challenging due to the unavailability of sufficient parallel data for
low-resource languages like Bengali. We provide Shironaam, a large-scale news headline generation dataset of a low-resource language
i.e., Bengali… See the full description on the dataset page: https://huggingface.co/datasets/dialect-ai/shironaam.arabic-dialects-gold20-code-switch
gold20-code-switch
Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic
lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20).
Each row embeds foreign material in a dialectal Arabic frame: inline
Latin-script English (and French, for the lects whose live contact language
is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi
(Latin-written Arabic with digit gutturals).
Columns (TSV, UTF-8, one file per lect):
id… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20-code-switch.portuguese-dialects-ipa-synthetic
portuguese-dialects-ipa-synthetic
920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties
(European regional, insular, Brazilian regional, African/Asian/border national norms,
medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects),
Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho,
and Galician-Portuguese. Each row carries two IPA columns with distinct provenance.
Schema
sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.Multi-Arabic-dialectsbasque_dialect_machine_translationArabic_Dialects
Dataset Card for Arabic Dialects
Dataset Summary
The Arabic Dialects dataset is a collection of text samples representing multiple spoken Arabic dialects alongside Modern Standard Arabic (MSA). It is designed to help train and evaluate natural language processing (NLP) models on dialect identification, text classification, and understanding regional linguistic variations.
Languages and Dialects Included
Egyptian (EGY)
Gulf (GLF)
Levantine (LEV)… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Arabic_Dialects.DialectGenIf you find our work helpful, please kindly cite our work :)
@article{zhou2025dialectgen,
title={DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation},
author={Zhou, Yu and An, Sohyun and Deng, Haikang and Yin, Da and Peng, Clark and Hsieh, Cho-Jui and Chang, Kai-Wei and Peng, Nanyun},
journal={arXiv preprint arXiv:2510.14949},
year={2025}
}
DialectDialect task as used in "Tokenization is Sensitive to Language Variation" paper, see arxiv.
@article{wegmann2025tokenization,
title={Tokenization is Sensitive to Language Variation},
author={Wegmann, Anna and Nguyen, Dong and Jurgens, David},
journal={arXiv preprint arXiv:2502.15343},
year={2025}
}
GLUE-dialectGLUE+dialect tasks used in "Tokenization is Sensitive to Language Variation paper", Arxiv link
@article{wegmann2025tokenization,
title={Tokenization is Sensitive to Language Variation},
author={Wegmann, Anna and Nguyen, Dong and Jurgens, David},
journal={arXiv preprint arXiv:2502.15343},
year={2025}
}
organic-gulf-arabic-dialect-dataset
Organic Gulf Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Gulf Arabic (Khaleeji) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-gulf-arabic-dialect-dataset.basque_dialect_identificationVulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words
If you use Vulgar Lexicon dataset, please cite the following paper:
@Article{app132111875,
AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito},
TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla},
JOURNAL = {Applied Sciences},
VOLUME = {13},
YEAR = {2023},
NUMBER = {21},
ARTICLE-NUMBER = {11875},
URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/kit-nlp/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.Arabic_dialects_to_MSAdialect-identificationorganic-levantine-arabic-dialect-dataset
Organic Levantine Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-levantine-arabic-dialect-dataset.arabic-dialect-text
Arabic Dialectal Text — gathered, lang-coded, IPA-enriched
A deduplicated collection of dialectal Arabic sentences assembled from openly-downloadable
sources, every line tagged with a BCP-47 lang code. Saudi Arabic is the focus, but all
labelled dialects are retained. Built as the text side of a Saudi TTS / phonemizer pipeline.
Files
all.tsv — the corpus: id<TAB>lang<TAB>source<TAB>text.
all.enriched.tsv — adds two phonetic columns:… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialect-text.basque_dialect_normalizationhin_dialect_classificationVulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words
If you use Vulgar Lexicon dataset, please cite the following paper:
@Article{app132111875,
AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito},
TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla},
JOURNAL = {Applied Sciences},
VOLUME = {13},
YEAR = {2023},
NUMBER = {21},
ARTICLE-NUMBER = {11875},
URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/TanjimKIT/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.organic-sudanese-arabic-dialect-dataset
Organic Sudanese Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Sudanese Arabic dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-sudanese-arabic-dialect-dataset.organic-iraqi-arabic-dialect-dataset
Organic Iraqi Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Iraqi Arabic dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately captures… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-iraqi-arabic-dialect-dataset.Open-ended_Questions_dialectal_data
Dataset Summary
A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences.
For more details, please check the paper
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection
Citation Information
@inproceedings{yamani-etal-2024-kind,
title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.organic-maghrebi-arabic-dialect-dataset
Organic Maghrebi Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Maghrebi Arabic (Darija) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-maghrebi-arabic-dialect-dataset.arabic_dialectseast_java_dialect_instruct
Complaints From The East Javanese Dialect community
This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments.
Cyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>>
If you use CBDCB dataset, please cite the following paper:
@article{mahmud2023cyberbullying,
title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art},
author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito},
journal={Information Processing \& Management},
volume={60},
number={5},
pages={103454},
year={2023},
publisher={Elsevier}
}
organic-egyptian-arabic-dialect-dataset
Organic Egyptian Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-egyptian-arabic-dialect-dataset.
