datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dialect-preferences
DiaLLM — Pooled Preference Dataset (Implicit Thread)
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
45,690 preference pairs, pooling all three variety-specific sets
(Australian,
Northern British,
Indian) without
variety targeting. Used for implicit-thread DPO training, where the three
varieties are pooled rather than targeted individually, preserving the
variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.Arabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets.
Dataset Card for Arabic_Dialect_Identification
Dataset Summary
We present QADI, an automatically collected dataset of tweets belonging to a wide range of
country-level Arabic dialects covering 18 different countries in the Middle East and North
Africa region. Our method for building this dataset relies on applying multiple filters to identify
users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.arabic-dialect-corpus
Arabic Dialect Corpus
A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata.
Dataset Statistics
Metric
Value
Total Records
127,180
Total Tokens
5,802,324
Average Tokens per Record
45.62
Dialect Categories
5
Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.Dialectal-Arabic-MMLU
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
Dataset Summary
Dialectal-Arabic-MMLU is a large-scale, human-translated for MMLU.
We extend MMLU-Redux into 5 major dialects: Syrian, Egyptian, Emirati, Saudi, and Moroccan.
This data covers 21K QA pairs across 32 academic and professional domains.
More details, please check our paper on DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Dialectal-Arabic-MMLU.DialectGenIf you find our work helpful, please kindly cite our work :)
@article{zhou2025dialectgen,
title={DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation},
author={Zhou, Yu and An, Sohyun and Deng, Haikang and Yin, Da and Peng, Clark and Hsieh, Cho-Jui and Chang, Kai-Wei and Peng, Nanyun},
journal={arXiv preprint arXiv:2510.14949},
year={2025}
}
GLUE-dialectGLUE+dialect tasks used in "Tokenization is Sensitive to Language Variation paper", Arxiv link
@article{wegmann2025tokenization,
title={Tokenization is Sensitive to Language Variation},
author={Wegmann, Anna and Nguyen, Dong and Jurgens, David},
journal={arXiv preprint arXiv:2502.15343},
year={2025}
}
arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.dialect_model_demoarabic_multi_dialect_dialogueexam_zh_multitopic_dialect_culture
exam_zh_multitopic_dialect_culture
This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge.
📚 Description
The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories:
🗣️ Regional Dialect Tests
These assess language understanding across major Chinese dialects and topolects:
Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.sada-eou-saudi-dialectDIALECT_COPACyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>>
If you use CBDCB dataset, please cite the following paper:
@article{mahmud2023cyberbullying,
title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art},
author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito},
journal={Information Processing \& Management},
volume={60},
number={5},
pages={103454},
year={2023},
publisher={Elsevier}
}
visual_accent_dialect_archiveSource: https://www.youtube.com/@visualaccent/videos
All rights belong to the original dataset creator.
VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects")
We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR).
This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.dialect_model_data
shanghai-binary dataset
Train/test splits for Shanghai vs Not-Shanghai binary classification.
Contents
data/train.parquet
data/test.parquet
Each row contains:
audio: float array (mono)
sampling_rate: 16000
dialect/label: label (Shanghai=1, else 0)
