datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vietnamese-diacritic-restoration-corpusI have downloaded it from Kaggle. I sincerely thank the author for making it available.
yoruba-diacritic-restoration-dataset
Yorùbá Diacritic Restoration Dataset
Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text.
Dataset Details
Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded)
Raw dataset: 8,365 rows, annotated with tone_pattern, harmony_class/harmony_breakdown, and focus_tag columns
Adapted dataset: 4,968 rows… See the full description on the dataset page: https://huggingface.co/datasets/Natashadonoh/yoruba-diacritic-restoration-dataset.lorde-language-model-tuning
