CoolFace
Datasetpublic

oddadmix/msa-omnivoice-tts-v1

MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.

sourceHugging Faceupdated 3mo agoView on Hugging Face
1likes289downloads
Dataset Card

MSA-OmniVoice-v1

Dataset Description

MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts.

It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.

Dataset Summary

  • —Language: Modern Standard Arabic (MSA)
  • —Hours: ~50
  • —Speakers: 1 (Synthetic)
  • —Transcripts: Fully diacritized (تشكيل)
  • —Speech: Synthetic (OmniVoice)

Dataset Structure

Each sample contains:

  • —audio – Audio waveform
  • —text – Fully diacritized Arabic transcript
  • —duration – Audio duration (seconds)

Supported Tasks

  • —Text-to-Speech (TTS)
  • —Automatic Speech Recognition (ASR)
  • —Speech Representation Learning
  • —Speech Alignment

Intended Uses

This dataset can be used for:

  • —Fine-tuning Arabic TTS models
  • —Training Arabic ASR models
  • —Speech-language research
  • —Synthetic data augmentation

Limitations

  • —Single synthetic speaker only.
  • —Does not provide speaker diversity.
  • —Synthetic speech may not fully reflect natural human speech.
  • —Small artifacts may exist in generated audio or transcripts.

Acknowledgements

Speech in this dataset was generated using Lahgtna. The dataset was curated and released by Oddadmix to support open Arabic speech AI research.

Citation

bibtex
@dataset{oddadmix_msa_omnivoice_v1,
  author = {Ahmed Wasfy},
  title = {MSA-OmniVoice-v1},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/datasets/oddadmix/MSA-OmniVoice-v1}
}