datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
central-kurdish-tts4all
TTS4All Central Kurdish Speech Dataset
Dataset Summary
The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish).
The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers.
The corpus was designed to support:
Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.arabic-tts-saudi-multi-speaker-xtts
Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦
This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format.
📌 Overview
Language: Arabic (Saudi Dialect)
Format: LJSpeech
Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.)
Speakers: Multi-speaker (Male & Female)
Audio Format: WAV (mono recommended)
Sample Rate: 22050 Hz (recommended)
📂 Structure
all_data/
│
├── wavs/
│ ├── sample_0.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.The_Arabic_News_speech_Corpus_Dataset
Arabic News Speech Corpus Dataset
This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics.
Dataset Details
Dataset Description
This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.Arabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.arabic-speech-dataset
Field
Value
License
cc-by-nc-nd-4.0
Task Categories
Automatic Speech Recognition
Language
Arabic (ar)
Tags
Arabic, Speech, Audio, Speech Recognition, Machine Learning
Size Category
1K < n < 10K
🎧 Arabic Speech Dataset
📘 Overview
The Arabic Speech Dataset is a high-quality speech audio dataset built for developing, training, and evaluating advanced AI voice systems. It provides 76 hours of audio data distributed across 558 files, available… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/arabic-speech-dataset.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/mohamedalgmaal/sada2022-arabic-tts.petra-sadouski-moi-shybolet-autabiiagrafichnyia-arabeski-petra-sadouski
Мой шыболет. Аўтабіяграфічныя арабэскі
Metadata
Author: Пётра Садоўскі
Title: Мой шыболет. Аўтабіяграфічныя арабэскі
Narrator: Пётра Садоўскі
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/petra-sadouski-moi-shybolet-autabiiagrafichnyia-arabeski-petra-sadouski.
