datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-whisper-multidialect-processedarabic-multidialect
Arabic Whisper Multi-Dialect ASR Dataset
A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning.
Dataset Description
This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks.
Dialects Included
Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts
Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/arabic-multidialect.Multi-Arabic-dialectsarabic-multidialect-emotional-speech-demo
DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech
A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request.
Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.kurdish-multidialect-asr-benchmark
Kurdish Dialect Speech Corpus
This project aims to provide a multi-dialect speech recognition benchmark for the Kurdish language. The Central Kurdish portion is the same as the Asosoft benchmark. The sentences were originally written in Central Kurdish (CKB), translated into other Kurdish dialects, and then recorded by native speakers.
The current version includes three Kurdish dialects: Central Kurdish, Northern Kurdish, and Southern Kurdish. A Hawrami version and the Badini… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/kurdish-multidialect-asr-benchmark.arabic-whisper-multidialect
Arabic Whisper Multi-Dialect ASR Dataset
A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning.
Dataset Description
This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks.
Dialects Included
Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts
Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/MadLook/arabic-whisper-multidialect.arabic_multi_dialect_dialoguearabic-whisper-multidialect-processed-small
Arabic Whisper Multi-Dialect - Processed (Small)
Dataset Description
This is a preprocessed version of the Arabic multi-dialect speech dataset, ready for fine-tuning OpenAI's Whisper models. The dataset contains audio features extracted and formatted specifically for Whisper training.
Size: 40% subset of the full arabic-whisper-multidialect dataset
Total Examples: 43,091 samples
Format: Pre-computed Whisper input features (mel spectrograms) and tokenized labels
Purpose:… See the full description on the dataset page: https://huggingface.co/datasets/MadLook/arabic-whisper-multidialect-processed-small.chained-sql-multidialectArabic-Multi-Dialect-Checkpoints
