datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
saudi1Lahgtna-saudi
Lahgtna Saudi (Mans1611/Lahgtna-saudi)
Saudi-dialect subset prepared for Arabic ASR fine-tuning
(from oddadmix/dialectal-arabic-lahgtna-v2, filtered to language == "sa").
Splits
Split
Rows
train
11,030
test
581
Columns
audio, text, language, duration
Features
{'audio': Audio(sampling_rate=16000, decode=False, num_channels=None, stream_index=None), 'text': Value('string'), 'language': Value('string'), 'duration':… See the full description on the dataset page: https://huggingface.co/datasets/Mans1611/Lahgtna-saudi.saudi2saudi-english-code-switching-datasetarabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.saudi_dialect_asrv1.0saudi2_consaudi-dialect-speech-female
🌍 Saudi Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
🗂️ Dataset Columns
Column
Description
audio
The audio chunk (22,050 Hz, mono WAV)
duration
Chunk duration in seconds
base_transcription
Transcript from the base Arabic ASR model
dialectal_transcription
Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.saudi1_con_tempsaudi-tts-synthetic-200k30k-SADA22_Saudisaudi-dialect-speech-male268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample
Description
Arabic(Saudi) Multi-stream Spontaneous Dialogue Smartphone speech dataset-Customer Service. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1627?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample.arabic-tts-saudi-multi-speaker-xtts
Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦
This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format.
📌 Overview
Language: Arabic (Saudi Dialect)
Format: LJSpeech
Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.)
Speakers: Multi-speaker (Male & Female)
Audio Format: WAV (mono recommended)
Sample Rate: 22050 Hz (recommended)
📂 Structure
all_data/
│
├── wavs/
│ ├── sample_0.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset
Description
사우디아라비아 아랍어(Arabic-Saudi) 멀티스트림 자연 대화 스마트폰 고객 서비스 음성 데이터셋입니다. 다양한 고객 서비스 상황에서 자유롭게 대화하는 방식으로 수집되었으며, 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 268명의 아랍어 원어민 화자로부터 데이터를 수집하여 다양한 화자 특성을 반영했으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1627?source=hf.kr
Specifications
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset.SaudiTalk
SaudiTalk: A Multi-Source Dialectal Speech Dataset from Saudi Arabia
Dataset Summary
SaudiTalk is a curated and human-verified Arabic speech dataset covering three major Saudi dialects: Hijazi, Ha’il, and Southern. The dataset is constructed from publicly available social media content and is designed to support research in Automatic Speech Recognition (ASR), dialect identification, and Arabic speech processing.
Key Features
3 Saudi dialects:… See the full description on the dataset page: https://huggingface.co/datasets/ranaRan689/SaudiTalk.SDAIANCAI-Saudilang-Code-Switch-CorpusSaudi_Podcasts_ASRyasser-aldosari-saudi-centersaudi_asrsaudi_arabic_accent60H-SADA22-Saudiarabic-tts-saudi-audio-datasetsaudi-tts-synthetic-100kFree_Dialogue_in_Saudi_Arabia_Corpus
Description
This dataset covers multiple scenarios such as banking, healthcare, insurance, sales, telecom, travel. The speakers are gender evenly, and each set of the audio is approximately 0.5 hour.
For more details, please refer to the link: https://dataoceanai.com/datasets/asr/free-dialogue-in-saudi-arabia-corpus/
Specification
ID:
King-ASR-919
SIZE:
113 hours
LANGUAGE:
Saudi Arabia
SPEAKERS:
100
AGES:
18-45 years old
DEVICES:
Mobile
saudi-podcast-1hrSaudi-data-ttssaudi-cs-dataset
