datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dhivehi-audios-82-spk
Dhivehi Synthetic Voice and Speech Augmentation Dataset
This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-82-spk.dhivehi-tts-preprocesseddhivehi-conversations-turn
Dhivehi Conversations (Turn-Based)
This is an experimental synthetic dataset of turn-based Dhivehi conversations created for testing and fine-tuning dialogue models, text-to-speech (TTS), and multi-turn speaker-aware systems.
This dataset is artificially constructed and not based on real conversations. It is intended for research experimentation only and may not always produce contextually accurate results.
Dataset Source
Derived from alakxender/voice-synthetic… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-conversations-turn.dhivehi-tts-female-refined-splitdhivehi-audio-kn
Dhivehi Audio Dataset
This is a Dhivehi (Maldivian) speech synthesis dataset with audio recordings, text transcriptions, and phonetic annotations. All recordings are by one speaker.
Dataset Overview
This dataset provides 4,170 high-quality audio samples in Dhivehi.
Key Statistics
Metric
Value
Total Audio Files
4,170
Total Duration
5.87 hours (352.0 minutes)
Average Clip Length
5.06 seconds
Total Words
30,068
Unique Phonemes
14190
Unique… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audio-kn.dhivehi-audio-casts
Dhivehi-Audio-Casts
Audio dataset with transcripts and speaker characteristics extracted from Qdrant database.
Dataset Summary
This dataset contains 3,300 audio recordings with corresponding transcripts and speaker characteristics including:
Audio: WAV format audio files
Sentence: Transcribed text in Dhivehi
Speaker Demographics: Age, gender probabilities (male/female/child)
Audio Characteristics: Arousal, dominance, valence scores
Audio Length: Duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audio-casts.dhivehi-audio-casts-processeddhivehi-shaafiu-speechDhivehi Shaafiu Speech is a single speaker Dhivehi speech dataset created by [Javaabu Pvt. Ltd.](https://javaabu.com).
The dataset contains around 16.5 hrs of text read by professional Maldivian narrator Muhammadh Shaafiu.
The text used for the recordings were text scrapped from various Maldivian news websites.dhivehi-javaabu-speech-parquetdhivehi-tts-female-01dhivehi_speech_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Repository: (https://huggingface.co/datasets/saajidha/dhivehi_speech_dataset)
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
Out-of-Scope Use
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/saajidha/dhivehi_speech_dataset.dhivehi-audio-1000-splitdhivehi-majlis-speechDhivehi Majlis Speech is a Dhivehi speech dataset created from data annotated by [Javaabu Pvt. Ltd.](https://javaabu.com).
The dataset contains around 10.5 hrs of speech collected from parliament sessions at The Peoples Majlis of Maldives (Maldivian Parliament) consisting of audio from different MPs from 6 different sessions.dhivehi-tts-male-02-splitdhivehi-khadheeja-speechDhivehi Khadheeja Speech is a single speaker Dhivehi speech dataset created by [Javaabu Pvt. Ltd.](https://javaabu.com).
The dataset contains around 20 hrs of text read by professional Maldivian narrator Khadheeja Faaz.
The text used for the recordings were text scrapped from various Maldivian news websites.dhivehi-shaafiu-speech-traincommon-voice-dhivehi-maledhivehi-audios-ds2
Dhivehi Audio Dataset 2
A quality-filtered Dhivehi speech dataset with three subsets (bronze, silver, gold),
each representing a progressively stricter quality threshold.
Subsets
Subset
MOS
CTC gc
WER
Train
Test
bronze
≥3.0
≥0.70
—
65,838
7,316
silver
≥3.0
≥0.70
≤0.30
43,010
4,779
gold
≥3.5
≥0.75
≤0.20
7,058
785
MOS — Mean Opinion Score (1–4), subjective listening quality rating
CTC gc — CTC forced-alignment geo-confidence (0–1), measures… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-ds2.dhivehi-audios-ds1
Dhivehi Audio Dataset 1
A quality-filtered Dhivehi speech dataset with three subsets (bronze, silver, gold),
each representing a progressively stricter quality threshold.
Subsets
Subset
MOS
CTC gc
WER
Train
Test
bronze
≥3.0
≥0.70
—
224,321
24,925
silver
≥3.0
≥0.70
≤0.30
131,389
14,599
gold
≥3.5
≥0.75
≤0.20
15,535
1,727
MOS — Mean Opinion Score (1–4), subjective listening quality rating
CTC gc — CTC forced-alignment geo-confidence (0–1)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-ds1.dhivehi-tts-male-03-splitdhivehi-tts-female-01-split
Dataset Card for "dhivehi-tts-female-01-split"
More Information needed
Quran-dhivehi-translation-audio-verse-by-versedhivehi-tts-male-refined-splitdhivehi-mms-v5-combineddhivehi-tts-combined-splitscaling_dhivehi_stt
