CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesteq /arabic-dialects-gold20 arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20.texttext-to-speechn<1K0 likes834 downloads2mo agoHugging Face02drelhaj /Arabic-Dialects Arabic Dialects Dataset (Bivalency & Code-Switching) The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties: EGY – Egyptian Arabic GLF – Gulf Arabic LAV – Levantine Arabic NOR – North African / Tunisian Arabic MSA – Modern Standard Arabic The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.texttext-classification10K<n<100K4 likes408 downloads10mo agoHugging Face03dialect-ai /shironaam Dataset Card for Shironaam Corpus Dataset Summary Automatic headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers. However, the performance of headline generation systems remains challenging due to the unavailability of sufficient parallel data for low-resource languages like Bengali. We provide Shironaam, a large-scale news headline generation dataset of a low-resource language i.e., Bengali… See the full description on the dataset page: https://huggingface.co/datasets/dialect-ai/shironaam.texttext-generation100K<n<1M6 likes192 downloads3y agoHugging Face04Salesteq /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes184 downloads2mo agoHugging Face05TigreGotico /portuguese-dialects-ipa-synthetic portuguese-dialects-ipa-synthetic 920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties (European regional, insular, Brazilian regional, African/Asian/border national norms, medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects), Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho, and Galician-Portuguese. Each row carries two IPA columns with distinct provenance. Schema sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.texttext-to-speechn<1K0 likes153 downloads2mo agoHugging Face06aminedjebbie /Multi-Arabic-dialectstext10K<n<100K1 likes123 downloads5y agoHugging Face07jaio98 /basque_dialect_machine_translationtext100K<n<1M0 likes86 downloads3mo agoHugging Face08fatymahaly /Arabic_Dialects Dataset Card for Arabic Dialects Dataset Summary The Arabic Dialects dataset is a collection of text samples representing multiple spoken Arabic dialects alongside Modern Standard Arabic (MSA). It is designed to help train and evaluate natural language processing (NLP) models on dialect identification, text classification, and understanding regional linguistic variations. Languages and Dialects Included Egyptian (EGY) Gulf (GLF) Levantine (LEV)… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Arabic_Dialects.texttext-classificationn<1K1 likes77 downloads14d agoHugging Face09uclanlp /DialectGenIf you find our work helpful, please kindly cite our work :) @article{zhou2025dialectgen, title={DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation}, author={Zhou, Yu and An, Sohyun and Deng, Haikang and Yin, Da and Peng, Clark and Hsieh, Cho-Jui and Chang, Kai-Wei and Peng, Nanyun}, journal={arXiv preprint arXiv:2510.14949}, year={2025} } tabular1K<n<10K2 likes55 downloads11mo agoHugging Face10AnnaWegmann /DialectDialect task as used in "Tokenization is Sensitive to Language Variation" paper, see arxiv. @article{wegmann2025tokenization, title={Tokenization is Sensitive to Language Variation}, author={Wegmann, Anna and Nguyen, Dong and Jurgens, David}, journal={arXiv preprint arXiv:2502.15343}, year={2025} } text10K<n<100K0 likes52 downloads1y agoHugging Face11AnnaWegmann /GLUE-dialectGLUE+dialect tasks used in "Tokenization is Sensitive to Language Variation paper", Arxiv link @article{wegmann2025tokenization, title={Tokenization is Sensitive to Language Variation}, author={Wegmann, Anna and Nguyen, Dong and Jurgens, David}, journal={arXiv preprint arXiv:2502.15343}, year={2025} } tabular100K<n<1M0 likes49 downloads1y agoHugging Face12ebubekr53 /organic-gulf-arabic-dialect-dataset Organic Gulf Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Gulf Arabic (Khaleeji) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-gulf-arabic-dialect-dataset.text1K<n<10K0 likes35 downloads3mo agoHugging Face13jaio98 /basque_dialect_identificationtext1K<n<10K0 likes32 downloads4mo agoHugging Face14kit-nlp /Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words If you use Vulgar Lexicon dataset, please cite the following paper: @Article{app132111875, AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito}, TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla}, JOURNAL = {Applied Sciences}, VOLUME = {13}, YEAR = {2023}, NUMBER = {21}, ARTICLE-NUMBER = {11875}, URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/kit-nlp/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.texttext-classification1K<n<10K1 likes30 downloads3y agoHugging Face15CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face16PRAli22 /Arabic_dialects_to_MSAtext100K<n<1M10 likes27 downloads3y agoHugging Face17ThinkAI-Morocco /dialect-identificationtext10K<n<100K1 likes24 downloads2y agoHugging Face18ebubekr53 /organic-levantine-arabic-dialect-dataset Organic Levantine Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-levantine-arabic-dialect-dataset.text1K<n<10K0 likes24 downloads3mo agoHugging Face19Salesteq /arabic-dialect-textgated Arabic Dialectal Text — gathered, lang-coded, IPA-enriched A deduplicated collection of dialectal Arabic sentences assembled from openly-downloadable sources, every line tagged with a BCP-47 lang code. Saudi Arabic is the focus, but all labelled dialects are retained. Built as the text side of a Saudi TTS / phonemizer pipeline. Files all.tsv — the corpus: id<TAB>lang<TAB>source<TAB>text. all.enriched.tsv — adds two phonetic columns:… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialect-text.texttext-to-speech100K<n<1M0 likes24 downloads29d agoHugging Face20jaio98 /basque_dialect_normalizationtext1K<n<10K0 likes23 downloads4mo agoHugging Face21mlexplorer008 /hin_dialect_classificationtext1K<n<10K0 likes22 downloads2y agoHugging Face22TanjimKIT /Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words If you use Vulgar Lexicon dataset, please cite the following paper: @Article{app132111875, AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito}, TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla}, JOURNAL = {Applied Sciences}, VOLUME = {13}, YEAR = {2023}, NUMBER = {21}, ARTICLE-NUMBER = {11875}, URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/TanjimKIT/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.texttext-classification1K<n<10K1 likes21 downloads3y agoHugging Face23ebubekr53 /organic-sudanese-arabic-dialect-dataset Organic Sudanese Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Sudanese Arabic dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-sudanese-arabic-dialect-dataset.textn<1K0 likes18 downloads2mo agoHugging Face24ebubekr53 /organic-iraqi-arabic-dialect-dataset Organic Iraqi Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Iraqi Arabic dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately captures… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-iraqi-arabic-dialect-dataset.text1K<n<10K0 likes17 downloads3mo agoHugging Face25KIND-Dataset /Open-ended_Questions_dialectal_data Dataset Summary A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences. For more details, please check the paper The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection Citation Information @inproceedings{yamani-etal-2024-kind, title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.textquestion-answeringn<1K0 likes16 downloads3y agoHugging Face26ebubekr53 /organic-maghrebi-arabic-dialect-dataset Organic Maghrebi Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Maghrebi Arabic (Darija) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-maghrebi-arabic-dialect-dataset.text1K<n<10K0 likes16 downloads3mo agoHugging Face27Hamma-16 /arabic_dialectstext10K<n<100K2 likes13 downloads1y agoHugging Face28yukebrillianth /east_java_dialect_instruct Complaints From The East Javanese Dialect community This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments. textquestion-answeringn<1K0 likes12 downloads1y agoHugging Face29TanjimKIT /Cyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBgatedPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>> If you use CBDCB dataset, please cite the following paper: @article{mahmud2023cyberbullying, title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art}, author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito}, journal={Information Processing \& Management}, volume={60}, number={5}, pages={103454}, year={2023}, publisher={Elsevier} } tabulartext-classification1K<n<10K2 likes11 downloads3y agoHugging Face30ebubekr53 /organic-egyptian-arabic-dialect-dataset Organic Egyptian Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-egyptian-arabic-dialect-dataset.text1K<n<10K1 likes11 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.