CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M29 likes4.4k downloads2mo agoHugging Face02MohamedRashad /MASC-Arabic MASC Arabic Dataset Card Dataset Summary MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels. The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.audioautomatic-speech-recognition100K<n<1M8 likes2.8k downloads6mo agoHugging Face03oddadmix /arabic-audio-collection-algerian-loubna-stories Loubna Stories Arabic Speech Dataset Dataset Summary The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.audiotext-to-speech10K<n<100K0 likes1.9k downloads3mo agoHugging Face04moaead /dialectal-arabic-voices Dialectal Arabic Voices An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps). 47,194 recordings · approximately 8,775.8 hours · 461.06 GB Column Description audio Original audio, embedded in the Parquet file transcript_text Empty for now; ASR transcripts will be added later language Dialect code: ps (Palestinian) source Original channel or account name Audio retains its original… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.audioautomatic-speech-recognition10K<n<100K0 likes1.7k downloads14m agoHugging Face05MohamedRashad /mgb2-arabic MGB-2: Arabic Multi-Dialect Broadcast Media Recognition Dataset Description Dataset Summary The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/mgb2-arabic.audioautomatic-speech-recognition100K<n<1M7 likes1.3k downloads9mo agoHugging Face06MohamedRashad /common-voice-18-arabic Dataset Card for Common Voice 18 – Arabic Edition Dataset Summary This dataset is an unofficial Arabic-only extraction of Mozilla Common Voice Corpus 18.0, prepared for Automatic Speech Recognition (ASR) research and development. It is derived from the original Common Voice 18 release and filtered to include Arabic (ar) speech data only, while preserving the original dataset structure, splits, and metadata fields. The dataset consists of validated, unvalidated, and… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/common-voice-18-arabic.audioautomatic-speech-recognition100K<n<1M5 likes483 downloads9mo agoHugging Face07xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes358 downloads2y agoHugging Face08oddadmix /arabic-audio-collection-algerian-kahwa-postcast Kahwa Postcast Arabic Speech Dataset Dataset Summary The Kahwa Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 110 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-kahwa-postcast.audiotext-to-speech10K<n<100K2 likes353 downloads3mo agoHugging Face09ahmed220v /mgb2-arabic MGB-2: Arabic Multi-Dialect Broadcast Media Recognition Dataset Description Dataset Summary The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/ahmed220v/mgb2-arabic.audioautomatic-speech-recognition100K<n<1M0 likes302 downloads7mo agoHugging Face10oddadmix /arabic-audio-collection-moroccan-ameed Ameed Moroccan Arabic Speech Dataset Dataset Summary The Ameed Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 176 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-ameed.audiotext-to-speech10K<n<100K0 likes295 downloads3mo agoHugging Face11MohamedRashad /MGB-3-Arabic Dataset Card for MGB-3 Arabic Speech Recognition Dataset Summary The MGB-3 Arabic dataset is a multi-genre collection of Egyptian Arabic speech extracted from YouTube videos, designed for speech recognition in challenging, real-world conditions. Unlike its predecessor MGB-2 which focused on broadcast TV news, MGB-3 emphasizes dialectal Arabic across diverse content types. The dataset contains approximately 16 hours of Egyptian Arabic speech from 80 YouTube videos… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MGB-3-Arabic.audioautomatic-speech-recognition1K<n<10K7 likes284 downloads9mo agoHugging Face12MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K34 likes260 downloads2y agoHugging Face13ismaeeelxd /Egyptian-Arabic-Lectures Egyptian Arabic Lectures Dataset The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts. Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.audioautomatic-speech-recognition1K<n<10K3 likes246 downloads3mo agoHugging Face14oddadmix /arabic-audio-collection-sudanese-sudan-podcast Sudan Podcast Arabic Speech Dataset Dataset Summary The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content. Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.audiotext-to-speech10K<n<100K1 likes212 downloads2mo agoHugging Face15oddadmix /arabic-audio-collection-mostafa-mahmoud Mostafa Mahmoud Arabic Speech Dataset Dataset Summary The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud. The dataset was created to support Arabic speech technology research and development, including: Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mostafa-mahmoud.audiotext-to-speech10K<n<100K13 likes206 downloads3mo agoHugging Face16oddadmix /arabic-audio-collection-mohamed-khairy Mohamed Khairy Arabic Speech Dataset Dataset Summary The Mohamed Khairy Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 430 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mohamed-khairy.audiotext-to-speech10K<n<100K7 likes203 downloads3mo agoHugging Face17HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes197 downloads5mo agoHugging Face18amine-khelif /arabic-multidialect Arabic Whisper Multi-Dialect ASR Dataset A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning. Dataset Description This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks. Dialects Included Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/arabic-multidialect.audioautomatic-speech-recognition100K<n<1M0 likes194 downloads8mo agoHugging Face19abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes159 downloads2mo agoHugging Face20oddadmix /arabic-audio-collection-algerian-rawi Rawi Postcast Arabic Speech Dataset Dataset Summary The Rawi Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 51 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-rawi.audiotext-to-speech1K<n<10K0 likes156 downloads3mo agoHugging Face21oddadmix /arabic-audio-collection-moroccan-wak3i Mak3i Moroccan Arabic Speech Dataset Dataset Summary The Mak3i Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 70 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-wak3i.audiotext-to-speech10K<n<100K0 likes156 downloads3mo agoHugging Face22tunis-ai /arabic_speech_corpus Dataset Card for Arabic Speech Corpus Dataset Summary This Speech corpus has been developed as part of PhD work carried out by Nawar Halabi at the University of Southampton. The corpus was recorded in south Levantine Arabic (Damascian accent) using a professional studio. Synthesized speech as an output using this corpus has produced a high quality, natural voice. Supported Tasks and Leaderboards [Needs More Information] Languages The audio is in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/arabic_speech_corpus.audioautomatic-speech-recognition1K<n<10K5 likes148 downloads2y agoHugging Face23Rabe3 /egyptian-arabic-tts-diacritized Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.tabulartext-to-speech10K<n<100K0 likes140 downloads1mo agoHugging Face24oddadmix /arabic-audio-collection-sudanese-nuuar Nuuar Sudanese Arabic Speech Dataset Dataset Summary The Nuuar Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 75 hours of speech recordings and corresponding transcripts. Sudanese Arabic remains one of the most underrepresented Arabic varieties in speech technology. This dataset directly addresses that gap by providing long-form, natural, dialectal Sudanese speech from a single consistent speaker, making… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-nuuar.audiotext-to-speech10K<n<100K0 likes127 downloads2mo agoHugging Face25NightPrince /Arabic-professional-voice Arabic Professional Voice A high-quality, single-speaker Arabic Text-to-Speech (TTS) dataset recorded by a professional speaker. All transcriptions include full Tashkeel (diacritical marks), making it directly suitable for training neural TTS systems without additional text normalization. Dataset Summary Property Value Language Arabic — Modern Standard Arabic (MSA) Utterances 439 Speaker 1 (professional male speaker) Sampling Rate 16 kHz Format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/Arabic-professional-voice.audiotext-to-speechn<1K2 likes126 downloads7mo agoHugging Face26oddadmix /arabic-audio-collection-syrian-podcast Syrian Postcast Arabic Speech Dataset Dataset Summary The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.audiotext-to-speech10K<n<100K2 likes125 downloads3mo agoHugging Face27datahiveai /arabic-multidialect-emotional-speech-demo DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request. Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.audioautomatic-speech-recognitionn<1K3 likes112 downloads5mo agoHugging Face28oddadmix /arabic-audio-collection-sudanese-ahmed-gobara Ahmed Gobara Sudanese Arabic Speech Dataset Dataset Summary The Ahmed Gobara Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 19 hours of speech recordings and corresponding transcripts. While compact, the dataset offers a clean, consistent single-speaker resource in Sudanese Arabic — an Arabic variety with very few open speech resources — making it especially valuable for voice cloning, speaker adaptation… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-ahmed-gobara.audiotext-to-speech1K<n<10K1 likes111 downloads2mo agoHugging Face29badrex /arabic-speech-SADA22-MSA Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.audioautomatic-speech-recognition1K<n<10K2 likes88 downloads1y agoHugging Face30FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes83 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.