datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dhivehi-noisy-sentences
Dhivehi Noisy Sentences Dataset
This dataset contains parallel examples of clean text and text with introduced errors across three categories: spelling, grammar, and punctuation.
Dataset Description
This dataset is designed to train models that can correct errors in Dhivehi text. Each example consists of:
clean_text: The correct, error-free Dhivehi text
noisy_text: The same text with introduced errors
error_type: The category of error (spelling, grammar, or punctuation)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-noisy-sentences.dhivehi-news-corpus
Thaana News Corpus Dataset
A comprehensive collection of news articles in Thaana script, extracted from various Maldivian news sources.
Data Format
Each record in the dataset contains:
title: The article title in Thaana script
content: The main article content in Thaana script
Dataset Updates
This dataset is regularly updated with new articles. Updates are performed incrementally, preserving existing data while adding new content.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-news-corpus.dhivehi-corpus
ދިވެހި Corpus — Dhivehi Text Corpus
Clean text corpus for the Dhivehi (Maldivian) language.
Built for NLP research and language model training.
Dataset Summary
Split
Docs
Tokens
Train
430,695
~81.6M
Validation
23,924
~4.5M
Test
23,924
~4.6M
Total
478,543
~90.6M
Language: Dhivehi (dv) — written in Thaana script (Unicode U+0780–U+07BF)
License: CC-BY-4.0
Avg quality score: 0.958 / 1.0
Duplicates: 0 (MinHash LSH deduplication at 80% threshold)… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-corpus.dhivehi-summaries
Dhivehi Text Summarization Dataset
Dataset Description
This dataset contains Dhivehi (Maldivian) text summarization pairs, consisting of original news articles and their corresponding summaries. The dataset is designed to support the development of Dhivehi text summarization models and advance NLP research for low-resource languages.
Dataset Summary
Language: Dhivehi (dv)
Task: Text Summarization
Domain: News Articles
Format: Original content paired with… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-summaries.dhivehi-transliteration-pairs
Dhivehi Transliteration Pairs
This dataset contains 187,908 aligned sentence pairs in Dhivehi (Thaana script) and its romanized (transliterated) Latin script form, making it a valuable resource for machine translation, cross-lingual NLP research, and bilingual corpus analysis.
Dataset Details
Language pair: English ↔ Dhivehi (Thaana script)
Train examples: 150,326
Test examples: 37,582
Total examples: 187,908
Dataset Structure
DatasetDict({
train:… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-transliteration-pairs.dhivehi-stories
Dhivehi Stories Dataset
A collection of 13,956 translated Dhivehi (Maldivian) stories with summaries and standardized metadata. Text is in Thaana script and cleaned for consistency, though sentence flow may not always be natural and some non-local names may appear.
Dataset Description
This dataset contains 13,956 stories written in Dhivehi (ދިވެހި), the official language of the Maldives. Each story includes:
Original story text in Dhivehi
Summary in Dhivehi
Metadata… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-stories.dhivehi-instruct
Dhivehi Instruct v1
A high-quality, multi-task instruction dataset designed to improve Natural
Language Processing (NLP) capabilities for the Dhivehi language. Formatted in a
standard conversational structure, it provides clean, contextually accurate
data for training, fine-tuning, and evaluating language models on
Dhivehi-specific tasks.
v1 contains no synthetic text. Every assistant output is human-written
corpus text or a deterministic, rule-based transform of it. No LLM… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-instruct.dhivehi-speeches
Dhivehi Speeches Dataset
This dataset contains speeches and articles scraped from the official website of the President's Office of the Maldives (https://presidency.gov.mv/). It is intended for research and language modeling purposes, especially for the Dhivehi language.
Dataset Structure
The dataset is provided as a single Parquet file: dhivehi_speeches.parquet.
Each row contains:
topic: The title or main topic of the speech/article.
speaker: The name of the speaker… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-speeches.dhivehi-QA
Dhivehi Question–Answer Instruction Dataset
This repository contains a simple JSON dataset of Dhivehi (ދިވެހި) question–answer pairs formatted as instruction–input–output triples.It is intended for instruction tuning, chatbot prototyping, and question-answering tasks in Dhivehi.
Dataset Overview
Language: Dhivehi (dv)
Script: Thaana
Format: JSON list of records
Fields:
instruction: prompt or task directive (e.g., “Answer the question in Dhivehi”)
input: the… See the full description on the dataset page: https://huggingface.co/datasets/axmeeabdhullo/dhivehi-QA.dhivehi-news-corpus
Thaana News Corpus Dataset
A comprehensive collection of news articles in Thaana script, extracted from various Maldivian news sources.
Data Format
Each record in the dataset contains:
title: The article title in Thaana script
content: The main article content in Thaana script
Dataset Updates
This dataset is regularly updated with new articles. Updates are performed incrementally, preserving existing data while adding new content.
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dhivehi-news-corpus.dhivehi-stories
Dhivehi Stories
A large collection of Dhivehi-language fiction stories scraped from esfiya.com, the most popular online fiction portal in the Maldives. The dataset contains 17,975 short stories, novelettes, and serial chapters, spanning from 2012 to early 2026.
Dataset contents
Field
Description
id
Unique story ID from esfiya.com
title
Title of the story/chapter (Thaana)
date
Publication date (UTC)
modified
Last modified date (UTC)
url
Source URL on… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-stories.dhivehi-legal-text-parallel
Dhivehi-English Legal Parallel Corpus
Dataset Description
A high-quality parallel corpus of 56,556 Dhivehi-English sentence pairs extracted from 200 Maldivian legal documents. This dataset is deduplicated and cleaned for machine translation and bilingual model training.
Dataset Summary
Languages: Dhivehi (dv) ↔ English (en)
Total Pairs: 56,556
Source Laws: 200
Duplicates Removed: 31,235
Average Dhivehi Length: 173.6 characters
Average English… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-legal-text-parallel.
