oddadmix/msa-omnivoice-tts-v1
MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.
MSA-OmniVoice-v1
Dataset Description
MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts.
It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.
Dataset Summary
- Language: Modern Standard Arabic (MSA)
- Hours: ~50
- Speakers: 1 (Synthetic)
- Transcripts: Fully diacritized (تشكيل)
- Speech: Synthetic (OmniVoice)
Dataset Structure
Each sample contains:
audio– Audio waveformtext– Fully diacritized Arabic transcriptduration– Audio duration (seconds)
Supported Tasks
- Text-to-Speech (TTS)
- Automatic Speech Recognition (ASR)
- Speech Representation Learning
- Speech Alignment
Intended Uses
This dataset can be used for:
- Fine-tuning Arabic TTS models
- Training Arabic ASR models
- Speech-language research
- Synthetic data augmentation
Limitations
- Single synthetic speaker only.
- Does not provide speaker diversity.
- Synthetic speech may not fully reflect natural human speech.
- Small artifacts may exist in generated audio or transcripts.
Acknowledgements
Speech in this dataset was generated using Lahgtna. The dataset was curated and released by Oddadmix to support open Arabic speech AI research.
Citation
@dataset{oddadmix_msa_omnivoice_v1,
author = {Ahmed Wasfy},
title = {MSA-OmniVoice-v1},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/oddadmix/MSA-OmniVoice-v1}
}