saajidha/dhivehi_speech_dataset
Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Repository: (https://huggingface.co/datasets/saajidha/dhivehi_speech_dataset) Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use Out-of-Scope Use Dataset… See the full description on the dataset page: https://huggingface.co/datasets/saajidha/dhivehi_speech_dataset.
Dataset Card for Dataset Name
<!-- Dhivehi Speech Recognition Dataset. -->
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
<!-- This dataset includes short audio clips of spoken Dhivehi along with their corresponding text transcriptions. The recordings were collected from native Dhivehi speakers under controlled conditions. Audio is in .wav format sampled at 16kHz, and transcriptions are in UTF-8 encoded text files.
- Curated by: by my self- saajidha mohamed
- Funded by [optional]: [More Information Needed]
- Shared by [optional]: [More Information Needed]
- Language(s) (NLP): Dhivehi (
dv) - License: BigScience OpenRAIL-M
Dataset Sources [optional]
<!-- Provide the basic links for the dataset. -->
- Repository: (https://huggingface.co/datasets/saajidha/dhivehispeechdataset)
- Paper [optional]: [More Information Needed]
- Demo [optional]: [More Information Needed]
Uses
<!-- this dataset is designed for training and evaluating Dhivehi ASR systems. . -->
Direct Use
<!-- This dataset is designed for training and evaluating Dhivehi ASR systems, including fine-tuning models like Whisper or Wav2Vec2 for Dhivehi.. -->
Out-of-Scope Use
<!-- This dataset is not intended for biometric identification or surveillance applications. Use must comply with the license. -->
Dataset Structure
<!-- Each example includes:
audio: The audio file (.wav)text: The transcript of the spoken audio
Splits (if available): train, validation, and test. -->
Dataset Creation
Curation Rationale
<!-- There is a lack of publicly available Dhivehi speech datasets. This dataset helps bridge that gap for ASR research in the Dhivehi language. -->
Source Data
<!-- sources include laws read, recorded and scripted for this purpose and public news read, recorded and scripted for this purpose, judgements read, recorded and scripted for this purpose.. -->
Data Collection and Processing
<!-- Audio was recorded by my self, a native speaker using mobile devices or microphones. The speech was transcribed manually and verified. Data was normalized to remove noise and silence, etc. -->
Who are the source data producers?
<!-- i by myself a Native speaker of Dhivehi from the Maldives. No personally identifiable information was collected or included. -->
Annotations [optional]
<!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. -->
Annotation process
<!-- Transcriptions were created by a native speaker and verified by a second annotator. -->
Who are the annotators?
<!-- The team consisted of linguists and volunteers fluent in Dhivehi. -->
Personal and Sensitive Information
<!-- No sensitive information is included. Data has been anonymized and filtered to avoid privacy risks. -->
Bias, Risks, and Limitations
<!-- The dataset may not cover all dialects or accents of Dhivehi. Urban speakers are more represented than rural ones. -->
Recommendations
<!-- Users should be cautious of dialectal bias when training models. The dataset may not generalize well to informal or spontaneous speech. -->
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
Citation [optional]
<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->
BibTeX:
@misc{saajidha2025dhivehi, title = {Dhivehi Speech Recognition Dataset}, author = {Saajidha Mohamed}, year = {2025}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/datasets/saajidha/dhivehispeechdataset}}, note = {Dataset for Dhivehi Automatic Speech Recognition (ASR)} APA:
Saajidha Mohamed. (2025). Dhivehi Speech Recognition Dataset [Dataset]. Hugging Face. https://huggingface.co/datasets/saajidha/dhivehispeechdataset
Glossary [optional]
<!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. -->
[More Information Needed]
More Information [optional]
[More Information Needed]
Dataset Card Authors [optional]
[More Information Needed]
Dataset Card Contact
Name: Saajidha Mohamed
Email: shiraan005@gmail.com
Hugging Face: https://huggingface.co/saajidha
