datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.dialectal-arabic-voices
Dialectal Arabic Voices
An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps).
50,652 recordings · approximately 9,268.1 hours · 486.11 GB
Column
Description
audio
Original audio, embedded in the Parquet file
transcript_text
Empty for now; ASR transcripts will be added later
language
Dialect code: ps (Palestinian)
source
Original channel or account name
Audio retains its original… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.
