Dialectal Arabic
dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.dialectal-arabic-voices
Dialectal Arabic Voices
An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps).
49,264 recordings · approximately 9,071.4 hours · 476.10 GB
Column
Description
audio
Original audio, embedded in the Parquet file
transcript_text
Empty for now; ASR transcripts will be added later
language
Dialect code: ps (Palestinian)
source
Original channel or account name
Audio retains its original… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.Dialectal-Arabic-MMLU
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
Dataset Summary
Dialectal-Arabic-MMLU is a large-scale, human-translated for MMLU.
We extend MMLU-Redux into 5 major dialects: Syrian, Egyptian, Emirati, Saudi, and Moroccan.
This data covers 21K QA pairs across 32 academic and professional domains.
More details, please check our paper on DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Dialectal-Arabic-MMLU.
