CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes80 downloads5mo agoHugging Face02IbrahimSalah /The_Arabic_News_speech_Corpus_Dataset Arabic News Speech Corpus Dataset This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics. Dataset Details Dataset Description This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.audioautomatic-speech-recognition1K<n<10K6 likes69 downloads2y agoHugging Face03FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes51 downloads5mo agoHugging Face04Speech-data /arabic-speech-dataset Field Value License cc-by-nc-nd-4.0 Task Categories Automatic Speech Recognition Language Arabic (ar) Tags Arabic, Speech, Audio, Speech Recognition, Machine Learning Size Category 1K < n < 10K 🎧 Arabic Speech Dataset 📘 Overview The Arabic Speech Dataset is a high-quality speech audio dataset built for developing, training, and evaluating advanced AI voice systems. It provides 76 hours of audio data distributed across 558 files, available… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/arabic-speech-dataset.audioautomatic-speech-recognitionn<1K0 likes34 downloads6mo agoHugging Face05Salesteq /arabic-dialect-textgated Arabic Dialectal Text — gathered, lang-coded, IPA-enriched A deduplicated collection of dialectal Arabic sentences assembled from openly-downloadable sources, every line tagged with a BCP-47 lang code. Saudi Arabic is the focus, but all labelled dialects are retained. Built as the text side of a Saudi TTS / phonemizer pipeline. Files all.tsv — the corpus: id<TAB>lang<TAB>source<TAB>text. all.enriched.tsv — adds two phonetic columns:… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialect-text.texttext-to-speech100K<n<1M0 likes19 downloads1mo agoHugging Face06driodnexus /droidnexus-arabic-editorial-speech-scorecard-mini DroidNexus Arabic Editorial Speech Scorecard Mini A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable. Why this exists This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.tabularautomatic-speech-recognitionn<1K0 likes12 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.