Dhivehi
Datasets
All datasets matching “Dhivehi”dhivehi-audios-82-spk
Dhivehi Synthetic Voice and Speech Augmentation Dataset
This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-82-spk.dhivehi-image-text
Dhivehi Image-Text Dataset
A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks.
Dataset Statistics
Total number of batches: 10
Total images across all batches: 394,212
Average images per batch: ~39,421
Split ratios:
Training: 80%
Validation: 10%
Test: 10%
Batch Details
Batch
Total Images
Train
Validation
Test
dv01-01
39659
31727
3966
3966
dv01-02
38989
31191
3899
3899
dv01-03
39360
31488
3936
3936… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-text.dhivehi-vrd-images
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Batch Statistics
Batch
Total Images
Train
Validation
Test
vrd-batch-1
474169
379335
47417
47417
vrd-batch-2
474493
379594
47449
47450
vrd-batch-3
475564
380451
47556… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-vrd-images.dhivehi-text-img-vqa
Dhivehi Image-Text Dataset
A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks.
Note: This dataset has been cleaned, and a 'question' field has been added from the alakxender/dhivehi-image-text collection.
dhivehi-image-bbox-prompt
Dhivehi Image Bounding Box Prompt Dataset
This dataset, alakxender/dhivehi-image-bbox-prompt, contains 58,738 images annotated with COCO-style bounding boxes and Dhivehi (Thaana script) text, along with layout categories such as Text, Title, Picture, Caption, and Columns. It is designed for OCR, document layout analysis, and multimodal vision–language research focused on Dhivehi.
Dataset
Each row includes:
image — the RGB image (preserved original dimensions)
width… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-prompt.dhivehi-image-bbox-ds-fmt
Dhivehi Image Bounding Box Dataset - DeepSeek Format
Dataset Description
This dataset is a transformed version of alakxender/dhivehi-image-bbox-prompt, specifically formatted to align with DeepSeek OCR model requirements for training vision-language models with grounding capabilities.
The original dataset contained Dhivehi (Thaana script) text with bounding box annotations. This version restructures the annotations into DeepSeek's grounding token format, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-ds-fmt.
