Natashadonoh/yoruba-diacritic-restoration-dataset
Yorùbá Diacritic Restoration Dataset Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text. Dataset Details Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded) Raw dataset: 8,365 rows, annotated with tone_pattern, harmony_class/harmony_breakdown, and focus_tag columns Adapted dataset: 4,968 rows… See the full description on the dataset page: https://huggingface.co/datasets/Natashadonoh/yoruba-diacritic-restoration-dataset.
014
