CoolFace
Datasetpublic

FatimahEmadEldin/alsanaa-emirati-arabic-asr

Alsanaa — Traditional Emirati Arabic Speech Dataset (ASR) A curated and preprocessed corpus for traditional Emirati (Gulf) Arabic Automatic Speech Recognition. It pairs short Emirati-dialect speech recordings with cleaned Arabic transcriptions, covering customs, etiquette, and oral tradition. Examples: 102 (train (102)) Approx. total audio: 4.54 hours Audio format on the Hub: decoded Audio feature at 16000 Hz Language: Arabic — traditional Emirati / Gulf dialect (ar)… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/alsanaa-emirati-arabic-asr.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes48downloads
Dataset Card

Alsanaa — Traditional Emirati Arabic Speech Dataset (ASR)

A curated and preprocessed corpus for traditional Emirati (Gulf) Arabic Automatic Speech Recognition. It pairs short Emirati-dialect speech recordings with cleaned Arabic transcriptions, covering customs, etiquette, and oral tradition.

  • —Examples: 102 (train (102))
  • —Approx. total audio: 4.54 hours
  • —Audio format on the Hub: decoded Audio feature at 16000 Hz
  • —Language: Arabic — traditional Emirati / Gulf dialect (ar)

⚠️ Attribution — please read

This dataset was created and curated by Maha AlBlooki, derived from recordings of the Aloula radio station and the Alsanaa book by Abdullah bin Dalmook. This Hub repository is a re-hosting of the original work published at <https://github.com/MahaAlBlooki/alsanaa-emirati-dataset> for convenient loading with the 🤗 datasets library. All credit belongs to the original author and the underlying sources. If you are the rights holder and want changes or removal, please open a discussion on this repo.

Dataset structure

Each example has:

fieldtypedescription
idstringutterance / file id (matches the source N.mp3)
audioAudiothe speech recording, decoded at 16000 Hz
transcriptionstringcleaned Arabic transcription of the utterance
duration_sfloatclip duration in seconds (from the source metadata.csv)
sourcestringprovenance note

Usage

python
from datasets import load_dataset

ds = load_dataset("FatimahEmadEldin/alsanaa-emirati-arabic-asr", split="train")
print(ds[0]["transcription"])
print(ds[0]["audio"]["sampling_rate"], ds[0]["audio"]["array"].shape)

Preprocessing (as described by the original author)

  • —Diacritics removal to normalise Arabic script (e.g. hamza variants → bare alif; tashkeel stripped).
  • —Punctuation removal for transcription consistency.
  • —Audio: silence/music cropping, resampling toward 16 kHz mono, volume normalisation.
Note: some diacritics remain in the source text; transcriptions are re-hosted faithfully, as-is.

Known caveats

  • —The source ships 103 audio files and 103 metadata rows, but 102 transcriptions. File `103` has audio + duration but no transcript. This build used the policy MISSING_TRANSCRIPT_POLICY = "drop".
  • —Recordings vary in length (≈ a few seconds to several minutes).

Intended uses

Traditional Emirati Arabic ASR training/evaluation, dialect identification, and phonological / sociolinguistic study of Gulf Arabic.

License

Published here under other. The original repository did not include an explicit data license; the audio derives from third-party sources (Aloula radio; the Alsanaa book). Confirm you have the right to use/redistribute before relying on it.

Citation

bibtex
@misc{alsanaa-emirati-dataset-2025,
  author       = {Maha AlBlooki},
  title        = {Traditional Emirati Arabic Speech Dataset: Curated from Aloula Radio and the Alsanaa Book},
  year         = {2025},
  howpublished = {\url{https://github.com/MahaAlBlooki/alsanaa-emirati-dataset}},
  note         = {Original dataset by Maha AlBlooki. Re-hosted on the Hugging Face Hub.}
}

Contact (original author)

📧 mahablooki@hotmail.com