CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01surrey-nlp /dialect-preferences DiaLLM — Pooled Preference Dataset (Implicit Thread) Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 45,690 preference pairs, pooling all three variety-specific sets (Australian, Northern British, Indian) without variety targeting. Used for implicit-thread DPO training, where the three varieties are pooled rather than targeted individually, preserving the variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.tabulartext-generation10K<n<100K0 likes233 downloads29d agoHugging Face02Abdelrahman-Rezk /Arabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets. Dataset Card for Arabic_Dialect_Identification Dataset Summary We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset relies on applying multiple filters to identify users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.tabular100K<n<1M12 likes165 downloads4y agoHugging Face03dataflare /arabic-dialect-corpus Arabic Dialect Corpus A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata. Dataset Statistics Metric Value Total Records 127,180 Total Tokens 5,802,324 Average Tokens per Record 45.62 Dialect Categories 5 Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.tabulartext-generation100K<n<1M1 likes158 downloads8mo agoHugging Face04MBZUAI /Dialectal-Arabic-MMLU DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models Dataset Summary Dialectal-Arabic-MMLU is a large-scale, human-translated for MMLU. We extend MMLU-Redux into 5 major dialects: Syrian, Egyptian, Emirati, Saudi, and Moroccan. This data covers 21K QA pairs across 32 academic and professional domains. More details, please check our paper on DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Dialectal-Arabic-MMLU.tabularmultiple-choice10K<n<100K1 likes152 downloads4mo agoHugging Face05uclanlp /DialectGenIf you find our work helpful, please kindly cite our work :) @article{zhou2025dialectgen, title={DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation}, author={Zhou, Yu and An, Sohyun and Deng, Haikang and Yin, Da and Peng, Clark and Hsieh, Cho-Jui and Chang, Kai-Wei and Peng, Nanyun}, journal={arXiv preprint arXiv:2510.14949}, year={2025} } tabular1K<n<10K2 likes55 downloads11mo agoHugging Face06AnnaWegmann /GLUE-dialectGLUE+dialect tasks used in "Tokenization is Sensitive to Language Variation paper", Arxiv link @article{wegmann2025tokenization, title={Tokenization is Sensitive to Language Variation}, author={Wegmann, Anna and Nguyen, Dong and Jurgens, David}, journal={arXiv preprint arXiv:2502.15343}, year={2025} } tabular100K<n<1M0 likes49 downloads1y agoHugging Face07CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face08karenlu653 /dialect_model_demotabularaudio-classificationn<1K0 likes23 downloads1y agoHugging Face09Kamyar-zeinalipour /arabic_multi_dialect_dialoguetabular10K<n<100K0 likes18 downloads1y agoHugging Face10somosnlp-hackathon-2025 /exam_zh_multitopic_dialect_culture exam_zh_multitopic_dialect_culture This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge. 📚 Description The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories: 🗣️ Regional Dialect Tests These assess language understanding across major Chinese dialects and topolects: Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.tabularmultiple-choicen<1K0 likes15 downloads1y agoHugging Face11mahmoudsaalama /sada-eou-saudi-dialecttabular1K<n<10K0 likes14 downloads10mo agoHugging Face12jjelinska /DIALECT_COPAtabular1K<n<10K0 likes12 downloads1y agoHugging Face13TanjimKIT /Cyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBgatedPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>> If you use CBDCB dataset, please cite the following paper: @article{mahmud2023cyberbullying, title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art}, author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito}, journal={Information Processing \& Management}, volume={60}, number={5}, pages={103454}, year={2023}, publisher={Elsevier} } tabulartext-classification1K<n<10K2 likes11 downloads3y agoHugging Face14Berkeley-NLP /visual_accent_dialect_archivegatedSource: https://www.youtube.com/@visualaccent/videos All rights belong to the original dataset creator. VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects") We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR). This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.audio1K<n<10K0 likes7 downloads7mo agoHugging Face15karenlu653 /dialect_model_data shanghai-binary dataset Train/test splits for Shanghai vs Not-Shanghai binary classification. Contents data/train.parquet data/test.parquet Each row contains: audio: float array (mono) sampling_rate: 16000 dialect/label: label (Shanghai=1, else 0) tabular1K<n<10K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.