CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesteq /arabic-dialects-gold20 arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20.texttext-to-speechn<1K0 likes834 downloads2mo agoHugging Face02drelhaj /Arabic-Dialects Arabic Dialects Dataset (Bivalency & Code-Switching) The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties: EGY – Egyptian Arabic GLF – Gulf Arabic LAV – Levantine Arabic NOR – North African / Tunisian Arabic MSA – Modern Standard Arabic The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.texttext-classification10K<n<100K4 likes408 downloads10mo agoHugging Face03Salesteq /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes184 downloads2mo agoHugging Face04aminedjebbie /Multi-Arabic-dialectstext10K<n<100K1 likes123 downloads5y agoHugging Face05fatymahaly /Arabic_Dialects Dataset Card for Arabic Dialects Dataset Summary The Arabic Dialects dataset is a collection of text samples representing multiple spoken Arabic dialects alongside Modern Standard Arabic (MSA). It is designed to help train and evaluate natural language processing (NLP) models on dialect identification, text classification, and understanding regional linguistic variations. Languages and Dialects Included Egyptian (EGY) Gulf (GLF) Levantine (LEV)… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Arabic_Dialects.texttext-classificationn<1K1 likes77 downloads14d agoHugging Face06ebubekr53 /organic-gulf-arabic-dialect-dataset Organic Gulf Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Gulf Arabic (Khaleeji) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-gulf-arabic-dialect-dataset.text1K<n<10K0 likes35 downloads3mo agoHugging Face07CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face08PRAli22 /Arabic_dialects_to_MSAtext100K<n<1M10 likes27 downloads3y agoHugging Face09ebubekr53 /organic-levantine-arabic-dialect-dataset Organic Levantine Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-levantine-arabic-dialect-dataset.text1K<n<10K0 likes24 downloads3mo agoHugging Face10Salesteq /arabic-dialect-textgated Arabic Dialectal Text — gathered, lang-coded, IPA-enriched A deduplicated collection of dialectal Arabic sentences assembled from openly-downloadable sources, every line tagged with a BCP-47 lang code. Saudi Arabic is the focus, but all labelled dialects are retained. Built as the text side of a Saudi TTS / phonemizer pipeline. Files all.tsv — the corpus: id<TAB>lang<TAB>source<TAB>text. all.enriched.tsv — adds two phonetic columns:… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialect-text.texttext-to-speech100K<n<1M0 likes24 downloads29d agoHugging Face11ebubekr53 /organic-sudanese-arabic-dialect-dataset Organic Sudanese Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Sudanese Arabic dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-sudanese-arabic-dialect-dataset.textn<1K0 likes18 downloads2mo agoHugging Face12ebubekr53 /organic-iraqi-arabic-dialect-dataset Organic Iraqi Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Iraqi Arabic dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately captures… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-iraqi-arabic-dialect-dataset.text1K<n<10K0 likes17 downloads3mo agoHugging Face13ebubekr53 /organic-maghrebi-arabic-dialect-dataset Organic Maghrebi Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Maghrebi Arabic (Darija) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-maghrebi-arabic-dialect-dataset.text1K<n<10K0 likes16 downloads3mo agoHugging Face14Hamma-16 /arabic_dialectstext10K<n<100K2 likes13 downloads1y agoHugging Face15ebubekr53 /organic-egyptian-arabic-dialect-dataset Organic Egyptian Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-egyptian-arabic-dialect-dataset.text1K<n<10K1 likes11 downloads3mo agoHugging Face16amany99 /arabic-dialect-to-msatext1K<n<10K0 likes7 downloads5mo agoHugging Face17Mayssm /rafeeq-arabic-medical-dialect Rafiq Arabic Medical Dialect Dataset Dataset summary Rafiq is an Arabic dataset for dialect-to-simple-Arabic normalization/translation of health-related expressions. The primary task is to convert a colloquial Arabic expression into simpler Arabic while preserving its meaning and avoiding added diagnoses or symptoms. medical_department is provided as an optional auxiliary text-classification/routing label. It is not the primary task and must not be interpreted as… See the full description on the dataset page: https://huggingface.co/datasets/Mayssm/rafeeq-arabic-medical-dialect.texttranslationn<1K0 likes4h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.