CoolFace
Datasetpublic

alakxender/dhivehi-audios-82-spk

Dhivehi Synthetic Voice and Speech Augmentation Dataset This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-82-spk.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes2kdownloads
Dataset Card

Dhivehi Synthetic Voice and Speech Augmentation Dataset

This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre variance.

Process

  • —Base text source: sentences from a Dhivehi news corpus.
  • —TTS model: speech model fine-tuned for Dhivehi phonetics.
  • —Voice cloning: reference recordings used to condition synthetic speakers.
  • —Augmentations:
  • —speed & tempo variation
  • —dynamic range compression
  • —pitch shifting (± semitones)
  • —formant warping & spectral noise
  • —reverb & background mix-in (random: 50-80 samples added to each subset)
  • —pronunciation drift simulation
  • —Generation time: ~36 hours of continuous synthesis.
  • —Sampling rate: 16 kHz PCM WAV.

Each row in the metadata includes:

FieldDescription
audioPath to .wav file
sentenceDhivehi text string
speaker_idOriginal speaker tag (e.g. fh_00)
subset_idMerged canonical speaker (e.g. f_00)
gendermale / female
presenceheavy / medium / light (augmentation intensity)

Dataset

SplitTotal SamplesDuration (hrs)SpeakersAvg Len (s)
Female (heavy)78 625153.657.04
Female (medium)235 862459.6157.02
Female (light)157 299306.8127.02
Male (heavy)77 639169.357.85
Male (medium)235 871512.1157.82
Male (light)471 7291 026.3307.83

Total: ≈ 1 257 024 samples (≈ 2 627 hours, 82 unique speakers)

Speaker Composition

  • —Male: ≈ 65 % of samples (45 speakers)
  • —Female: ≈ 35 % of samples (37 speakers)
  • —Per-speaker duration: ~30 – 34 hours (balanced)
  • —Pronunciation depth: each speaker appears under multiple augmentation presets (heavy, medium, light), producing distinct acoustic conditions.

Text Statistics

  • —Average sentence length: ~118 characters
  • —Median length: ~112 characters
  • —Range: 8 – 692 characters (max after token cleanup)
  • —Texts cover news, politics, society, and general narration topics.

Intended Use

Designed primarily for:

  • —Fine-tuning and evaluation of Dhivehi TTS systems
  • —Automatic Speech Recognition (ASR) robustness tests
  • —Speaker embedding and voice transfer experiments
  • —Cross-speaker adaptation and data-augmentation research

Not recommended for direct human listening or production-grade speech models without validation.

Disclaimer

This dataset contains synthetic audio generated for research and evaluation purposes only. It does not represent real human voices, nor does it reflect any individual’s identity or opinion.