CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aranemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.2k downloads3mo agoHugging Face02oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M29 likes3.7k downloads2mo agoHugging Face03MohamedRashad /MASC-Arabic MASC Arabic Dataset Card Dataset Summary MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels. The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.audioautomatic-speech-recognition100K<n<1M8 likes2.4k downloads6mo agoHugging Face04oddadmix /arabic-audio-collection-algerian-loubna-stories Loubna Stories Arabic Speech Dataset Dataset Summary The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.audiotext-to-speech10K<n<100K0 likes1.9k downloads3mo agoHugging Face05MohamedRashad /mgb2-arabic MGB-2: Arabic Multi-Dialect Broadcast Media Recognition Dataset Description Dataset Summary The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/mgb2-arabic.audioautomatic-speech-recognition100K<n<1M7 likes1.3k downloads9mo agoHugging Face06moaead /dialectal-arabic-voices Dialectal Arabic Voices An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps). 44,589 recordings · approximately 8,372.7 hours · 441.03 GB Column Description audio Original audio, embedded in the Parquet file transcript_text Empty for now; ASR transcripts will be added later language Dialect code: ps (Palestinian) source Original channel or account name Audio retains its original… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.audioautomatic-speech-recognition10K<n<100K0 likes877 downloads35m agoHugging Face07MohamedRashad /common-voice-18-arabic Dataset Card for Common Voice 18 – Arabic Edition Dataset Summary This dataset is an unofficial Arabic-only extraction of Mozilla Common Voice Corpus 18.0, prepared for Automatic Speech Recognition (ASR) research and development. It is derived from the original Common Voice 18 release and filtered to include Arabic (ar) speech data only, while preserving the original dataset structure, splits, and metadata fields. The dataset consists of validated, unvalidated, and… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/common-voice-18-arabic.audioautomatic-speech-recognition100K<n<1M5 likes467 downloads9mo agoHugging Face08xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes357 downloads2y agoHugging Face09oddadmix /arabic-audio-collection-algerian-kahwa-postcast Kahwa Postcast Arabic Speech Dataset Dataset Summary The Kahwa Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 110 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-kahwa-postcast.audiotext-to-speech10K<n<100K2 likes341 downloads3mo agoHugging Face10ahmed220v /mgb2-arabic MGB-2: Arabic Multi-Dialect Broadcast Media Recognition Dataset Description Dataset Summary The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/ahmed220v/mgb2-arabic.audioautomatic-speech-recognition100K<n<1M0 likes304 downloads7mo agoHugging Face11MohamedRashad /MGB-3-Arabic Dataset Card for MGB-3 Arabic Speech Recognition Dataset Summary The MGB-3 Arabic dataset is a multi-genre collection of Egyptian Arabic speech extracted from YouTube videos, designed for speech recognition in challenging, real-world conditions. Unlike its predecessor MGB-2 which focused on broadcast TV news, MGB-3 emphasizes dialectal Arabic across diverse content types. The dataset contains approximately 16 hours of Egyptian Arabic speech from 80 YouTube videos… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MGB-3-Arabic.audioautomatic-speech-recognition1K<n<10K7 likes283 downloads9mo agoHugging Face12oddadmix /arabic-audio-collection-moroccan-ameed Ameed Moroccan Arabic Speech Dataset Dataset Summary The Ameed Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 176 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-ameed.audiotext-to-speech10K<n<100K0 likes276 downloads3mo agoHugging Face13MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K34 likes247 downloads2y agoHugging Face14oddadmix /arabic-audio-collection-mostafa-mahmoud Mostafa Mahmoud Arabic Speech Dataset Dataset Summary The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud. The dataset was created to support Arabic speech technology research and development, including: Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mostafa-mahmoud.audiotext-to-speech10K<n<100K13 likes236 downloads3mo agoHugging Face15ismaeeelxd /Egyptian-Arabic-Lectures Egyptian Arabic Lectures Dataset The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts. Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.audioautomatic-speech-recognition1K<n<10K3 likes234 downloads3mo agoHugging Face16aranemini /northern-kurdish-pseudolabel Northern Kurdish Raw Audio Collection Dataset Summary This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources. The corpus was assembled to support research and development in: Automatic Speech Recognition (ASR) Speech Translation (ST) Text-to-Speech (TTS) Self-Supervised Learning (SSL) Spoken Language Understanding (SLU) Low-Resource Speech Processing The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.audioautomatic-speech-recognition100K<n<1M1 likes234 downloads3mo agoHugging Face17oddadmix /arabic-audio-collection-mohamed-khairy Mohamed Khairy Arabic Speech Dataset Dataset Summary The Mohamed Khairy Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 430 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mohamed-khairy.audiotext-to-speech10K<n<100K7 likes222 downloads3mo agoHugging Face18amine-khelif /arabic-multidialect Arabic Whisper Multi-Dialect ASR Dataset A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning. Dataset Description This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks. Dialects Included Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/arabic-multidialect.audioautomatic-speech-recognition100K<n<1M0 likes194 downloads8mo agoHugging Face19HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes186 downloads5mo agoHugging Face20oddadmix /arabic-audio-collection-sudanese-sudan-podcast Sudan Podcast Arabic Speech Dataset Dataset Summary The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content. Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.audiotext-to-speech10K<n<100K1 likes184 downloads2mo agoHugging Face21abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes164 downloads2mo agoHugging Face22oddadmix /arabic-audio-collection-moroccan-wak3i Mak3i Moroccan Arabic Speech Dataset Dataset Summary The Mak3i Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 70 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-wak3i.audiotext-to-speech10K<n<100K0 likes151 downloads3mo agoHugging Face23aranemini /central-kurdish-tts4all TTS4All Central Kurdish Speech Dataset Dataset Summary The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish). The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers. The corpus was designed to support: Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.audiotext-to-speech10K<n<100K3 likes148 downloads3mo agoHugging Face24oddadmix /arabic-audio-collection-algerian-rawi Rawi Postcast Arabic Speech Dataset Dataset Summary The Rawi Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 51 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-rawi.audiotext-to-speech1K<n<10K0 likes146 downloads3mo agoHugging Face25tunis-ai /arabic_speech_corpus Dataset Card for Arabic Speech Corpus Dataset Summary This Speech corpus has been developed as part of PhD work carried out by Nawar Halabi at the University of Southampton. The corpus was recorded in south Levantine Arabic (Damascian accent) using a professional studio. Synthesized speech as an output using this corpus has produced a high quality, natural voice. Supported Tasks and Leaderboards [Needs More Information] Languages The audio is in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/arabic_speech_corpus.audioautomatic-speech-recognition1K<n<10K5 likes141 downloads2y agoHugging Face26Aratako /Magpie-Speech-Orpheus-125k Magpie-Speech-Orpheus-125k A ~125k-sample synthetic speech dataset generated by applying the Magpie instruction-synthesis approach to the Orpheus-TTS LLM-based text-to-speech model, then decoding audio tokens with the SNAC 24 kHz codec. Blog (EN): https://huggingface.co/blog/Aratako/magpie-speech Blog (JA): https://zenn.dev/aratako_lm/articles/87d8988d44ba4d This dataset is entirely synthetic: text prompts and audio tokens were produced by Orpheus-TTS and decoded to waveforms via… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Speech-Orpheus-125k.audiotext-to-speech100K<n<1M11 likes141 downloads1y agoHugging Face27Rabe3 /egyptian-arabic-tts-diacritized Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.tabulartext-to-speech10K<n<100K0 likes133 downloads1mo agoHugging Face28oddadmix /arabic-audio-collection-syrian-podcast Syrian Postcast Arabic Speech Dataset Dataset Summary The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.audiotext-to-speech10K<n<100K2 likes122 downloads3mo agoHugging Face29NightPrince /Arabic-professional-voice Arabic Professional Voice A high-quality, single-speaker Arabic Text-to-Speech (TTS) dataset recorded by a professional speaker. All transcriptions include full Tashkeel (diacritical marks), making it directly suitable for training neural TTS systems without additional text normalization. Dataset Summary Property Value Language Arabic — Modern Standard Arabic (MSA) Utterances 439 Speaker 1 (professional male speaker) Sampling Rate 16 kHz Format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/Arabic-professional-voice.audiotext-to-speechn<1K2 likes119 downloads7mo agoHugging Face30datahiveai /arabic-multidialect-emotional-speech-demo DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request. Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.audioautomatic-speech-recognitionn<1K3 likes112 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.