FatimahEmadEldin/alsanaa-emirati-arabic-asr
Alsanaa — Traditional Emirati Arabic Speech Dataset (ASR) A curated and preprocessed corpus for traditional Emirati (Gulf) Arabic Automatic Speech Recognition. It pairs short Emirati-dialect speech recordings with cleaned Arabic transcriptions, covering customs, etiquette, and oral tradition. Examples: 102 (train (102)) Approx. total audio: 4.54 hours Audio format on the Hub: decoded Audio feature at 16000 Hz Language: Arabic — traditional Emirati / Gulf dialect (ar)… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/alsanaa-emirati-arabic-asr.
Alsanaa — Traditional Emirati Arabic Speech Dataset (ASR)
A curated and preprocessed corpus for traditional Emirati (Gulf) Arabic Automatic Speech Recognition. It pairs short Emirati-dialect speech recordings with cleaned Arabic transcriptions, covering customs, etiquette, and oral tradition.
- Examples: 102 (train (102))
- Approx. total audio: 4.54 hours
- Audio format on the Hub: decoded
Audiofeature at 16000 Hz - Language: Arabic — traditional Emirati / Gulf dialect (
ar)
⚠️ Attribution — please read
This dataset was created and curated by Maha AlBlooki, derived from recordings of the Aloula radio station and the Alsanaa book by Abdullah bin Dalmook. This Hub repository is a re-hosting of the original work published at <https://github.com/MahaAlBlooki/alsanaa-emirati-dataset> for convenient loading with the 🤗 datasets library. All credit belongs to the original author and the underlying sources. If you are the rights holder and want changes or removal, please open a discussion on this repo.
Dataset structure
Each example has:
Usage
from datasets import load_dataset
ds = load_dataset("FatimahEmadEldin/alsanaa-emirati-arabic-asr", split="train")
print(ds[0]["transcription"])
print(ds[0]["audio"]["sampling_rate"], ds[0]["audio"]["array"].shape)Preprocessing (as described by the original author)
- Diacritics removal to normalise Arabic script (e.g. hamza variants → bare alif; tashkeel stripped).
- Punctuation removal for transcription consistency.
- Audio: silence/music cropping, resampling toward 16 kHz mono, volume normalisation.
Note: some diacritics remain in the source text; transcriptions are re-hosted faithfully, as-is.
Known caveats
- The source ships 103 audio files and 103 metadata rows, but 102 transcriptions. File `103` has audio + duration but no transcript. This build used the policy
MISSING_TRANSCRIPT_POLICY = "drop". - Recordings vary in length (≈ a few seconds to several minutes).
Intended uses
Traditional Emirati Arabic ASR training/evaluation, dialect identification, and phonological / sociolinguistic study of Gulf Arabic.
License
Published here under other. The original repository did not include an explicit data license; the audio derives from third-party sources (Aloula radio; the Alsanaa book). Confirm you have the right to use/redistribute before relying on it.
Citation
@misc{alsanaa-emirati-dataset-2025,
author = {Maha AlBlooki},
title = {Traditional Emirati Arabic Speech Dataset: Curated from Aloula Radio and the Alsanaa Book},
year = {2025},
howpublished = {\url{https://github.com/MahaAlBlooki/alsanaa-emirati-dataset}},
note = {Original dataset by Maha AlBlooki. Re-hosted on the Hugging Face Hub.}
}Contact (original author)
📧 mahablooki@hotmail.com
