datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dhivehi-audios-82-spk
Dhivehi Synthetic Voice and Speech Augmentation Dataset
This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-82-spk.dhivehi-image-text
Dhivehi Image-Text Dataset
A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks.
Dataset Statistics
Total number of batches: 10
Total images across all batches: 394,212
Average images per batch: ~39,421
Split ratios:
Training: 80%
Validation: 10%
Test: 10%
Batch Details
Batch
Total Images
Train
Validation
Test
dv01-01
39659
31727
3966
3966
dv01-02
38989
31191
3899
3899
dv01-03
39360
31488
3936
3936… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-text.dhivehi-vrd-images
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Batch Statistics
Batch
Total Images
Train
Validation
Test
vrd-batch-1
474169
379335
47417
47417
vrd-batch-2
474493
379594
47449
47450
vrd-batch-3
475564
380451
47556… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-vrd-images.dhivehi-text-img-vqa
Dhivehi Image-Text Dataset
A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks.
Note: This dataset has been cleaned, and a 'question' field has been added from the alakxender/dhivehi-image-text collection.
dhivehi-image-bbox-prompt
Dhivehi Image Bounding Box Prompt Dataset
This dataset, alakxender/dhivehi-image-bbox-prompt, contains 58,738 images annotated with COCO-style bounding boxes and Dhivehi (Thaana script) text, along with layout categories such as Text, Title, Picture, Caption, and Columns. It is designed for OCR, document layout analysis, and multimodal vision–language research focused on Dhivehi.
Dataset
Each row includes:
image — the RGB image (preserved original dimensions)
width… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-prompt.dhivehi-image-bbox-ds-fmt
Dhivehi Image Bounding Box Dataset - DeepSeek Format
Dataset Description
This dataset is a transformed version of alakxender/dhivehi-image-bbox-prompt, specifically formatted to align with DeepSeek OCR model requirements for training vision-language models with grounding capabilities.
The original dataset contained Dhivehi (Thaana script) text with bounding box annotations. This version restructures the annotations into DeepSeek's grounding token format, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-ds-fmt.dhivehi-noisy-sentences
Dhivehi Noisy Sentences Dataset
This dataset contains parallel examples of clean text and text with introduced errors across three categories: spelling, grammar, and punctuation.
Dataset Description
This dataset is designed to train models that can correct errors in Dhivehi text. Each example consists of:
clean_text: The correct, error-free Dhivehi text
noisy_text: The same text with introduced errors
error_type: The category of error (spelling, grammar, or punctuation)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-noisy-sentences.dhivehi-news-corpus
Thaana News Corpus Dataset
A comprehensive collection of news articles in Thaana script, extracted from various Maldivian news sources.
Data Format
Each record in the dataset contains:
title: The article title in Thaana script
content: The main article content in Thaana script
Dataset Updates
This dataset is regularly updated with new articles. Updates are performed incrementally, preserving existing data while adding new content.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-news-corpus.dhivehi-vrd-batch-1-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-tts-preprocesseddhivehiDhivehi dataset for MNT
dhivehi-vrd-batch-3-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-conversations-turn
Dhivehi Conversations (Turn-Based)
This is an experimental synthetic dataset of turn-based Dhivehi conversations created for testing and fine-tuning dialogue models, text-to-speech (TTS), and multi-turn speaker-aware systems.
This dataset is artificially constructed and not based on real conversations. It is intended for research experimentation only and may not always produce contextually accurate results.
Dataset Source
Derived from alakxender/voice-synthetic… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-conversations-turn.dhivehi-vrd-batch-6-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
Dhivehi-Englishdhivehi-img-txtsen
Dhivehi News Layout Dataset
This dataset contains generated page layouts for Dhivehi news content in various traditional publication styles, intended for document layout generation and analysis tasks.
Dataset Description
Dataset Summary
The dataset consists of Dhivehi news articles rendered in different traditional layout styles, with varying fonts and orientations. Each sample includes both the rendered image and the original text content.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-img-txtsen.dhivehi-vrd-batch-2-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-vrd-batch-5-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-synthetic-v1dhivehi-vrd-batch-4-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-tts-female-refined-splitdhivehi-layout-syn-lg-florence
Synthetic Dhivehi Document Layout Analysis Dataset
Overview
This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-lg-florence.dhivehi-english-parallel
Dhivehi-English Cleaned Parallel Corpus
This dataset contains parallel sentence pairs between English and Dhivehi (Thaana script). It was built by merging and cleaning multiple sources, ensuring only valid pairs (not null) were retained.
Dataset Summary
Languages: English (en) → Dhivehi (dv)
Total Examples: 484133
Average Length:
English: ~355 characters
Dhivehi: ~487 characters
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-parallel.dhivehi-corpus
ދިވެހި Corpus — Dhivehi Text Corpus
Clean text corpus for the Dhivehi (Maldivian) language.
Built for NLP research and language model training.
Dataset Summary
Split
Docs
Tokens
Train
430,695
~81.6M
Validation
23,924
~4.5M
Test
23,924
~4.6M
Total
478,543
~90.6M
Language: Dhivehi (dv) — written in Thaana script (Unicode U+0780–U+07BF)
License: CC-BY-4.0
Avg quality score: 0.958 / 1.0
Duplicates: 0 (MinHash LSH deduplication at 80% threshold)… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-corpus.dhivehi-audio-kn
Dhivehi Audio Dataset
This is a Dhivehi (Maldivian) speech synthesis dataset with audio recordings, text transcriptions, and phonetic annotations. All recordings are by one speaker.
Dataset Overview
This dataset provides 4,170 high-quality audio samples in Dhivehi.
Key Statistics
Metric
Value
Total Audio Files
4,170
Total Duration
5.87 hours (352.0 minutes)
Average Clip Length
5.06 seconds
Total Words
30,068
Unique Phonemes
14190
Unique… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audio-kn.dhivehi-layout-syn-lg-paligemma
Synthetic Dhivehi Document Layout Analysis Dataset
Overview
This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-lg-paligemma.dhivehi-english-word-translations
Word-Level Dhivehi-English Translation
Dataset Description
This dataset contains Dhivehi words along with their English translations from the models:
Google Gemini 2.5 Flash Lite (g_word)
Anthropic Claude 3.7 Sonnet (c_word)
Dataset Structure
dv_word: The original Dhivehi word
g_word: Translation by Google Gemini 2.5 Flash Lite
c_word: Translation by Anthropic Claude 3.7 Sonnet
best_word: Binary indicator of which translation is better (0 = Gemini 2.5 Flash… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-word-translations.dhivehi-vrd-batch-2-img-questions
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc.
Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Note: This dataset is a subset from alakxender/dhivehi-vrd-images.
dhivehi-english-translations
Dhivehi-English Translation Dataset
Dataset Description
This dataset contains 91,759 Dhivehi-English translation pairs extracted from news articles and other content. The dataset is designed for machine translation research, cross-lingual information retrieval, and Dhivehi language processing tasks.
Languages
Source: Dhivehi (dv) - The official language of the Maldives
Target: English (en)
Dataset Structure
Data Instances
{
"dhivehi":… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-translations.dhivehi-layout-syn-b1-paligemma
Synthetic Dhivehi Document Layout Analysis Dataset
Overview
This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-b1-paligemma.
