datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
diacritics-ro
Romanian Diacritic Restoration Dataset (dexonline)
Data splits for Romanian automatic diacritic restoration (ADR), derived from the dexonline lexicographic corpus.
Dataset Description
Each example contains a input (stripped text without diacritics), a target (gold-standard diacritized text), and an id.
Source: dexonline.ro -- open-access Romanian lexicographic resource
Training pairs: 299,633 examples (NFKD diacritic stripping + cedilla normalization)
Validation: 33,292… See the full description on the dataset page: https://huggingface.co/datasets/klusai/diacritics-ro.arabic-sentences-diacritics
Arabic Diacritization Dataset (70 ePub Collection)
This dataset consists of sentence pairs extracted from 70 Arabic ePub books sourced from the "أولو العلم" Telegram channel.
It is designed for training models focused on Arabic diacritization, text cleaning, or character-level language modeling.
Dataset Summary
The project started as a prototype with 6,464 sentence pairs from 11 books. In this second iteration, the dataset has been expanded to 108,778 sentence pairs from… See the full description on the dataset page: https://huggingface.co/datasets/0xAgamy/arabic-sentences-diacritics.
