datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.central-kurdish-pseudolabel
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus
Dataset Summary
This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).
The dataset was automatically generated using a pipeline composed of:
Speech segmentation
Automatic Speech Recognition (ASR)
Machine Translation (MT)
The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.MASC-Arabic
MASC Arabic Dataset Card
Dataset Summary
MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels.
The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition.
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.arabic-audio-collection-algerian-loubna-stories
Loubna Stories Arabic Speech Dataset
Dataset Summary
The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.dialectal-arabic-voices
Dialectal Arabic Voices
An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps).
49,264 recordings · approximately 9,071.4 hours · 476.10 GB
Column
Description
audio
Original audio, embedded in the Parquet file
transcript_text
Empty for now; ASR transcripts will be added later
language
Dialect code: ps (Palestinian)
source
Original channel or account name
Audio retains its original… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.mgb2-arabic
MGB-2: Arabic Multi-Dialect Broadcast Media Recognition
Dataset Description
Dataset Summary
The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/mgb2-arabic.northern-kurdish-raw-audio
Northern Kurdish Raw Audio Collection
Overview
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The collection was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-supervised Learning (SSL)
Spoken Language Understanding (SLU)
The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.common-voice-18-arabic
Dataset Card for Common Voice 18 – Arabic Edition
Dataset Summary
This dataset is an unofficial Arabic-only extraction of Mozilla Common Voice Corpus 18.0, prepared for Automatic Speech Recognition (ASR) research and development.
It is derived from the original Common Voice 18 release and filtered to include Arabic (ar) speech data only, while preserving the original dataset structure, splits, and metadata fields. The dataset consists of validated, unvalidated, and… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/common-voice-18-arabic.commonvoice-12.0-arabic-voice-converted
Dataset Card for Voice Converted Arabic Common Voice 12.0
This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.arabic-audio-collection-algerian-kahwa-postcast
Kahwa Postcast Arabic Speech Dataset
Dataset Summary
The Kahwa Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 110 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-kahwa-postcast.mgb2-arabic
MGB-2: Arabic Multi-Dialect Broadcast Media Recognition
Dataset Description
Dataset Summary
The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/ahmed220v/mgb2-arabic.arabic-audio-collection-moroccan-ameed
Ameed Moroccan Arabic Speech Dataset
Dataset Summary
The Ameed Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 176 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-ameed.MGB-3-Arabic
Dataset Card for MGB-3 Arabic Speech Recognition
Dataset Summary
The MGB-3 Arabic dataset is a multi-genre collection of Egyptian Arabic speech extracted from YouTube videos, designed for speech recognition in challenging, real-world conditions. Unlike its predecessor MGB-2 which focused on broadcast TV news, MGB-3 emphasizes dialectal Arabic across diverse content types.
The dataset contains approximately 16 hours of Egyptian Arabic speech from 80 YouTube videos… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MGB-3-Arabic.arabic-english-code-switching
Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨
The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning.
Citation
If you use this dataset, please cite it as follows:
@misc{rashad2024arabic,
author = {Mohamed Rashad},
title = {arabic-english-code-switching},
year = {2024},
publisher = {Hugging Face},
url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.northern-kurdish-pseudolabel
Northern Kurdish Raw Audio Collection
Dataset Summary
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The corpus was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-Supervised Learning (SSL)
Spoken Language Understanding (SLU)
Low-Resource Speech Processing
The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.arabic-audio-collection-sudanese-sudan-podcast
Sudan Podcast Arabic Speech Dataset
Dataset Summary
The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content.
Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.arabic-audio-collection-mostafa-mahmoud
Mostafa Mahmoud Arabic Speech Dataset
Dataset Summary
The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud.
The dataset was created to support Arabic speech technology research and development, including:
Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mostafa-mahmoud.arabic-audio-collection-mohamed-khairy
Mohamed Khairy Arabic Speech Dataset
Dataset Summary
The Mohamed Khairy Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 430 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mohamed-khairy.arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.southern-kurdish-raw-audio
Southern Kurdish Raw Audio Collection
Overview
This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech collected from publicly available media sources, podcasts, interviews, news broadcasts, and online programs.
The main sources are Aryen TV and Kurd Channel.
The recordings mainly consist of spontaneous and semi-spontaneous speech, covering diverse speakers, topics, and acoustic conditions. While the majority of the content is… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/southern-kurdish-raw-audio.arabic-multidialect
Arabic Whisper Multi-Dialect ASR Dataset
A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning.
Dataset Description
This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks.
Dialects Included
Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts
Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/arabic-multidialect.central-kurdish-audiobook-raw
Central Kurdish Audiobook Raw Audio Collection
Overview
This repository contains a large collection of raw Central Kurdish (Sorani Kurdish) audiobook recordings gathered from publicly available online sources.
The collection was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-supervised learning
The dataset contains approximately 4,300 hours of speech collected from 1026… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-audiobook-raw.arabic-english-code-switching-synthetic-asr
Synthetic Arabic-English Code-Switched Speech for ASR
This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository.
Configurations
Configuration
Train
Test
Publication status
synthetic
8,655
962
Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.arabic-audio-collection-algerian-rawi
Rawi Postcast Arabic Speech Dataset
Dataset Summary
The Rawi Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 51 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-rawi.arabic-audio-collection-moroccan-wak3i
Mak3i Moroccan Arabic Speech Dataset
Dataset Summary
The Mak3i Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 70 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-wak3i.arabic_speech_corpus
Dataset Card for Arabic Speech Corpus
Dataset Summary
This Speech corpus has been developed as part of PhD work carried out by Nawar Halabi at the University of Southampton. The corpus was recorded in south Levantine Arabic (Damascian accent) using a professional studio. Synthesized speech as an output using this corpus has produced a high quality, natural voice.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/arabic_speech_corpus.central-kurdish-tts4all
TTS4All Central Kurdish Speech Dataset
Dataset Summary
The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish).
The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers.
The corpus was designed to support:
Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.Magpie-Speech-Orpheus-125k
Magpie-Speech-Orpheus-125k
A ~125k-sample synthetic speech dataset generated by applying the Magpie instruction-synthesis approach to the Orpheus-TTS LLM-based text-to-speech model, then decoding audio tokens with the SNAC 24 kHz codec.
Blog (EN): https://huggingface.co/blog/Aratako/magpie-speech
Blog (JA): https://zenn.dev/aratako_lm/articles/87d8988d44ba4d
This dataset is entirely synthetic: text prompts and audio tokens were produced by Orpheus-TTS and decoded to waveforms via… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Speech-Orpheus-125k.arabic-audio-collection-moroccan-noone-stories
Noone Stories Moroccan Arabic Speech Dataset
Dataset Summary
The Noone Stories Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 156 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-noone-stories.
