CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alakxender /dhivehi-noisy-sentences Dhivehi Noisy Sentences Dataset This dataset contains parallel examples of clean text and text with introduced errors across three categories: spelling, grammar, and punctuation. Dataset Description This dataset is designed to train models that can correct errors in Dhivehi text. Each example consists of: clean_text: The correct, error-free Dhivehi text noisy_text: The same text with introduced errors error_type: The category of error (spelling, grammar, or punctuation)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-noisy-sentences.texttranslation1M<n<10M0 likes211 downloads10mo agoHugging Face02alakxender /dhivehi-news-corpus Thaana News Corpus Dataset A comprehensive collection of news articles in Thaana script, extracted from various Maldivian news sources. Data Format Each record in the dataset contains: title: The article title in Thaana script content: The main article content in Thaana script Dataset Updates This dataset is regularly updated with new articles. Updates are performed incrementally, preserving existing data while adding new content.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-news-corpus.texttranslation100K<n<1M0 likes183 downloads29d agoHugging Face03d3b4g /dhivehi-corpus ދިވެހި Corpus — Dhivehi Text Corpus Clean text corpus for the Dhivehi (Maldivian) language. Built for NLP research and language model training. Dataset Summary Split Docs Tokens Train 430,695 ~81.6M Validation 23,924 ~4.5M Test 23,924 ~4.6M Total 478,543 ~90.6M Language: Dhivehi (dv) — written in Thaana script (Unicode U+0780–U+07BF) License: CC-BY-4.0 Avg quality score: 0.958 / 1.0 Duplicates: 0 (MinHash LSH deduplication at 80% threshold)… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-corpus.tabulartext-generation100K<n<1M0 likes74 downloads6mo agoHugging Face04alakxender /dhivehi-summaries Dhivehi Text Summarization Dataset Dataset Description This dataset contains Dhivehi (Maldivian) text summarization pairs, consisting of original news articles and their corresponding summaries. The dataset is designed to support the development of Dhivehi text summarization models and advance NLP research for low-resource languages. Dataset Summary Language: Dhivehi (dv) Task: Text Summarization Domain: News Articles Format: Original content paired with… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-summaries.textsummarization10K<n<100K0 likes35 downloads1y agoHugging Face05alakxender /dhivehi-transliteration-pairs Dhivehi Transliteration Pairs This dataset contains 187,908 aligned sentence pairs in Dhivehi (Thaana script) and its romanized (transliterated) Latin script form, making it a valuable resource for machine translation, cross-lingual NLP research, and bilingual corpus analysis. Dataset Details Language pair: English ↔ Dhivehi (Thaana script) Train examples: 150,326 Test examples: 37,582 Total examples: 187,908 Dataset Structure DatasetDict({ train:… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-transliteration-pairs.texttranslation100K<n<1M0 likes33 downloads1y agoHugging Face06alakxender /dhivehi-stories Dhivehi Stories Dataset A collection of 13,956 translated Dhivehi (Maldivian) stories with summaries and standardized metadata. Text is in Thaana script and cleaned for consistency, though sentence flow may not always be natural and some non-local names may appear. Dataset Description This dataset contains 13,956 stories written in Dhivehi (ދިވެހި), the official language of the Maldives. Each story includes: Original story text in Dhivehi Summary in Dhivehi Metadata… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-stories.textfeature-extraction10K<n<100K1 likes23 downloads1y agoHugging Face07d3b4g /dhivehi-instruct Dhivehi Instruct v1 A high-quality, multi-task instruction dataset designed to improve Natural Language Processing (NLP) capabilities for the Dhivehi language. Formatted in a standard conversational structure, it provides clean, contextually accurate data for training, fine-tuning, and evaluating language models on Dhivehi-specific tasks. v1 contains no synthetic text. Every assistant output is human-written corpus text or a deterministic, rule-based transform of it. No LLM… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-instruct.texttext-generation10K<n<100K0 likes21 downloads3mo agoHugging Face08alakxender /dhivehi-speeches Dhivehi Speeches Dataset This dataset contains speeches and articles scraped from the official website of the President's Office of the Maldives (https://presidency.gov.mv/). It is intended for research and language modeling purposes, especially for the Dhivehi language. Dataset Structure The dataset is provided as a single Parquet file: dhivehi_speeches.parquet. Each row contains: topic: The title or main topic of the speech/article. speaker: The name of the speaker… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-speeches.texttext-generationn<1K0 likes20 downloads1y agoHugging Face09axmeeabdhullo /dhivehi-QA Dhivehi Question–Answer Instruction Dataset This repository contains a simple JSON dataset of Dhivehi (ދިވެހި) question–answer pairs formatted as instruction–input–output triples.It is intended for instruction tuning, chatbot prototyping, and question-answering tasks in Dhivehi. Dataset Overview Language: Dhivehi (dv) Script: Thaana Format: JSON list of records Fields: instruction: prompt or task directive (e.g., “Answer the question in Dhivehi”) input: the… See the full description on the dataset page: https://huggingface.co/datasets/axmeeabdhullo/dhivehi-QA.texttext-generationn<1K0 likes16 downloads1y agoHugging Face10mashey /dhivehi-news-corpus Thaana News Corpus Dataset A comprehensive collection of news articles in Thaana script, extracted from various Maldivian news sources. Data Format Each record in the dataset contains: title: The article title in Thaana script content: The main article content in Thaana script Dataset Updates This dataset is regularly updated with new articles. Updates are performed incrementally, preserving existing data while adding new content. Usage This… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dhivehi-news-corpus.texttranslation10K<n<100K0 likes15 downloads4mo agoHugging Face11d3b4g /dhivehi-stories Dhivehi Stories A large collection of Dhivehi-language fiction stories scraped from esfiya.com, the most popular online fiction portal in the Maldives. The dataset contains 17,975 short stories, novelettes, and serial chapters, spanning from 2012 to early 2026. Dataset contents Field Description id Unique story ID from esfiya.com title Title of the story/chapter (Thaana) date Publication date (UTC) modified Last modified date (UTC) url Source URL on… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/dhivehi-stories.tabulartext-generation10K<n<100K0 likes14 downloads6mo agoHugging Face12alakxender /dhivehi-legal-text-parallelgated Dhivehi-English Legal Parallel Corpus Dataset Description A high-quality parallel corpus of 56,556 Dhivehi-English sentence pairs extracted from 200 Maldivian legal documents. This dataset is deduplicated and cleaned for machine translation and bilingual model training. Dataset Summary Languages: Dhivehi (dv) ↔ English (en) Total Pairs: 56,556 Source Laws: 200 Duplicates Removed: 31,235 Average Dhivehi Length: 173.6 characters Average English… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-legal-text-parallel.tabulartranslation10K<n<100K0 likes12 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.