CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alakxender /dhivehi-audios-82-spk Dhivehi Synthetic Voice and Speech Augmentation Dataset This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-82-spk.audioautomatic-speech-recognition1M<n<10M2 likes2.2k downloads11mo agoHugging Face02alakxender /dhivehi-image-text Dhivehi Image-Text Dataset A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks. Dataset Statistics Total number of batches: 10 Total images across all batches: 394,212 Average images per batch: ~39,421 Split ratios: Training: 80% Validation: 10% Test: 10% Batch Details Batch Total Images Train Validation Test dv01-01 39659 31727 3966 3966 dv01-02 38989 31191 3899 3899 dv01-03 39360 31488 3936 3936… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-text.imagetext-to-image100K<n<1M0 likes1.2k downloads2y agoHugging Face03alakxender /dhivehi-vrd-images Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Batch Statistics Batch Total Images Train Validation Test vrd-batch-1 474169 379335 47417 47417 vrd-batch-2 474493 379594 47449 47450 vrd-batch-3 475564 380451 47556… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-vrd-images.imagevisual-question-answering1M<n<10M0 likes898 downloads1y agoHugging Face04alakxender /dhivehi-text-img-vqa Dhivehi Image-Text Dataset A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks. Note: This dataset has been cleaned, and a 'question' field has been added from the alakxender/dhivehi-image-text collection. imagevisual-question-answering100K<n<1M0 likes754 downloads1y agoHugging Face05alakxender /dhivehi-image-bbox-prompt Dhivehi Image Bounding Box Prompt Dataset This dataset, alakxender/dhivehi-image-bbox-prompt, contains 58,738 images annotated with COCO-style bounding boxes and Dhivehi (Thaana script) text, along with layout categories such as Text, Title, Picture, Caption, and Columns. It is designed for OCR, document layout analysis, and multimodal vision–language research focused on Dhivehi. Dataset Each row includes: image — the RGB image (preserved original dimensions) width… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-prompt.imageobject-detection10K<n<100K0 likes529 downloads1y agoHugging Face06alakxender /dhivehi-image-bbox-ds-fmt Dhivehi Image Bounding Box Dataset - DeepSeek Format Dataset Description This dataset is a transformed version of alakxender/dhivehi-image-bbox-prompt, specifically formatted to align with DeepSeek OCR model requirements for training vision-language models with grounding capabilities. The original dataset contained Dhivehi (Thaana script) text with bounding box annotations. This version restructures the annotations into DeepSeek's grounding token format, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-ds-fmt.imageimage-to-text10K<n<100K0 likes350 downloads10mo agoHugging Face07alakxender /dhivehi-noisy-sentences Dhivehi Noisy Sentences Dataset This dataset contains parallel examples of clean text and text with introduced errors across three categories: spelling, grammar, and punctuation. Dataset Description This dataset is designed to train models that can correct errors in Dhivehi text. Each example consists of: clean_text: The correct, error-free Dhivehi text noisy_text: The same text with introduced errors error_type: The category of error (spelling, grammar, or punctuation)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-noisy-sentences.texttranslation1M<n<10M0 likes205 downloads10mo agoHugging Face08alakxender /dhivehi-news-corpus Thaana News Corpus Dataset A comprehensive collection of news articles in Thaana script, extracted from various Maldivian news sources. Data Format Each record in the dataset contains: title: The article title in Thaana script content: The main article content in Thaana script Dataset Updates This dataset is regularly updated with new articles. Updates are performed incrementally, preserving existing data while adding new content.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-news-corpus.texttranslation100K<n<1M0 likes188 downloads28d agoHugging Face09alakxender /dhivehi-vrd-batch-1-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes165 downloads1y agoHugging Face10Serialtechlab /dhivehi-tts-preprocessedaudio10K<n<100K0 likes152 downloads7mo agoHugging Face11chopey /dhivehiDhivehi dataset for MNT text100K<n<1M0 likes148 downloads5y agoHugging Face12alakxender /dhivehi-vrd-batch-3-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes143 downloads1y agoHugging Face13alakxender /dhivehi-conversations-turn Dhivehi Conversations (Turn-Based) This is an experimental synthetic dataset of turn-based Dhivehi conversations created for testing and fine-tuning dialogue models, text-to-speech (TTS), and multi-turn speaker-aware systems. This dataset is artificially constructed and not based on real conversations. It is intended for research experimentation only and may not always produce contextually accurate results. Dataset Source Derived from alakxender/voice-synthetic… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-conversations-turn.audiotext-to-speech10K<n<100K0 likes136 downloads1y agoHugging Face14alakxender /dhivehi-vrd-batch-6-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes131 downloads1y agoHugging Face15MKK /Dhivehi-English0 likes121 downloads5y agoHugging Face16alakxender /dhivehi-img-txtsen Dhivehi News Layout Dataset This dataset contains generated page layouts for Dhivehi news content in various traditional publication styles, intended for document layout generation and analysis tasks. Dataset Description Dataset Summary The dataset consists of Dhivehi news articles rendered in different traditional layout styles, with varying fonts and orientations. Each sample includes both the rendered image and the original text content. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-img-txtsen.image10K<n<100K0 likes115 downloads2y agoHugging Face17alakxender /dhivehi-vrd-batch-2-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes95 downloads1y agoHugging Face18alakxender /dhivehi-vrd-batch-5-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes89 downloads1y agoHugging Face19mohamedrayyan /dhivehi-synthetic-v1textn<1K0 likes89 downloads4mo agoHugging Face20alakxender /dhivehi-vrd-batch-4-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes82 downloads1y agoHugging Face21Serialtechlab /dhivehi-tts-female-refined-splitaudio10K<n<100K0 likes80 downloads7mo agoHugging Face22alakxender /dhivehi-layout-syn-lg-florence Synthetic Dhivehi Document Layout Analysis Dataset Overview This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-lg-florence.image10K<n<100K0 likes76 downloads1y agoHugging Face23alakxender /dhivehi-english-parallel Dhivehi-English Cleaned Parallel Corpus This dataset contains parallel sentence pairs between English and Dhivehi (Thaana script). It was built by merging and cleaning multiple sources, ensuring only valid pairs (not null) were retained. Dataset Summary Languages: English (en) → Dhivehi (dv) Total Examples: 484133 Average Length: English: ~355 characters Dhivehi: ~487 characters Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-parallel.texttranslation100K<n<1M0 likes73 downloads1y agoHugging Face24d3b4g /dhivehi-corpus ދިވެހި Corpus — Dhivehi Text Corpus Clean text corpus for the Dhivehi (Maldivian) language. Built for NLP research and language model training. Dataset Summary Split Docs Tokens Train 430,695 ~81.6M Validation 23,924 ~4.5M Test 23,924 ~4.6M Total 478,543 ~90.6M Language: Dhivehi (dv) — written in Thaana script (Unicode U+0780–U+07BF) License: CC-BY-4.0 Avg quality score: 0.958 / 1.0 Duplicates: 0 (MinHash LSH deduplication at 80% threshold)… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-corpus.tabulartext-generation100K<n<1M0 likes72 downloads6mo agoHugging Face25alakxender /dhivehi-audio-kn Dhivehi Audio Dataset This is a Dhivehi (Maldivian) speech synthesis dataset with audio recordings, text transcriptions, and phonetic annotations. All recordings are by one speaker. Dataset Overview This dataset provides 4,170 high-quality audio samples in Dhivehi. Key Statistics Metric Value Total Audio Files 4,170 Total Duration 5.87 hours (352.0 minutes) Average Clip Length 5.06 seconds Total Words 30,068 Unique Phonemes 14190 Unique… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audio-kn.audiotext-to-speech1K<n<10K0 likes61 downloads5mo agoHugging Face26alakxender /dhivehi-layout-syn-lg-paligemma Synthetic Dhivehi Document Layout Analysis Dataset Overview This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-lg-paligemma.image10K<n<100K0 likes54 downloads1y agoHugging Face27alakxender /dhivehi-english-word-translations Word-Level Dhivehi-English Translation Dataset Description This dataset contains Dhivehi words along with their English translations from the models: Google Gemini 2.5 Flash Lite (g_word) Anthropic Claude 3.7 Sonnet (c_word) Dataset Structure dv_word: The original Dhivehi word g_word: Translation by Google Gemini 2.5 Flash Lite c_word: Translation by Anthropic Claude 3.7 Sonnet best_word: Binary indicator of which translation is better (0 = Gemini 2.5 Flash… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-word-translations.texttranslation100K<n<1M0 likes54 downloads1y agoHugging Face28mashey /dhivehi-vrd-batch-2-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes53 downloads4mo agoHugging Face29alakxender /dhivehi-english-translations Dhivehi-English Translation Dataset Dataset Description This dataset contains 91,759 Dhivehi-English translation pairs extracted from news articles and other content. The dataset is designed for machine translation research, cross-lingual information retrieval, and Dhivehi language processing tasks. Languages Source: Dhivehi (dv) - The official language of the Maldives Target: English (en) Dataset Structure Data Instances { "dhivehi":… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-translations.texttranslation10K<n<100K0 likes47 downloads1y agoHugging Face30alakxender /dhivehi-layout-syn-b1-paligemma Synthetic Dhivehi Document Layout Analysis Dataset Overview This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-b1-paligemma.image1K<n<10K0 likes42 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.