CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uclanlp /DialectGenIf you find our work helpful, please kindly cite our work :) @article{zhou2025dialectgen, title={DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation}, author={Zhou, Yu and An, Sohyun and Deng, Haikang and Yin, Da and Peng, Clark and Hsieh, Cho-Jui and Chang, Kai-Wei and Peng, Nanyun}, journal={arXiv preprint arXiv:2510.14949}, year={2025} } tabular1K<n<10K2 likes55 downloads11mo agoHugging Face02AnnaWegmann /GLUE-dialectGLUE+dialect tasks used in "Tokenization is Sensitive to Language Variation paper", Arxiv link @article{wegmann2025tokenization, title={Tokenization is Sensitive to Language Variation}, author={Wegmann, Anna and Nguyen, Dong and Jurgens, David}, journal={arXiv preprint arXiv:2502.15343}, year={2025} } tabular100K<n<1M0 likes49 downloads1y agoHugging Face03CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face04TanjimKIT /Cyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBgatedPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>> If you use CBDCB dataset, please cite the following paper: @article{mahmud2023cyberbullying, title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art}, author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito}, journal={Information Processing \& Management}, volume={60}, number={5}, pages={103454}, year={2023}, publisher={Elsevier} } tabulartext-classification1K<n<10K2 likes11 downloads3y agoHugging Face05Berkeley-NLP /visual_accent_dialect_archivegatedSource: https://www.youtube.com/@visualaccent/videos All rights belong to the original dataset creator. VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects") We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR). This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.audio1K<n<10K0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.