datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DialectGenIf you find our work helpful, please kindly cite our work :)
@article{zhou2025dialectgen,
title={DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation},
author={Zhou, Yu and An, Sohyun and Deng, Haikang and Yin, Da and Peng, Clark and Hsieh, Cho-Jui and Chang, Kai-Wei and Peng, Nanyun},
journal={arXiv preprint arXiv:2510.14949},
year={2025}
}
GLUE-dialectGLUE+dialect tasks used in "Tokenization is Sensitive to Language Variation paper", Arxiv link
@article{wegmann2025tokenization,
title={Tokenization is Sensitive to Language Variation},
author={Wegmann, Anna and Nguyen, Dong and Jurgens, David},
journal={arXiv preprint arXiv:2502.15343},
year={2025}
}
arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.Cyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>>
If you use CBDCB dataset, please cite the following paper:
@article{mahmud2023cyberbullying,
title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art},
author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito},
journal={Information Processing \& Management},
volume={60},
number={5},
pages={103454},
year={2023},
publisher={Elsevier}
}
visual_accent_dialect_archiveSource: https://www.youtube.com/@visualaccent/videos
All rights belong to the original dataset creator.
VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects")
We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR).
This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.
