CoolFace
Datasetpublic

Natashadonoh/yoruba-diacritic-restoration-dataset

Yorùbá Diacritic Restoration Dataset Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text. Dataset Details Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded) Raw dataset: 8,365 rows, annotated with tone_pattern, harmony_class/harmony_breakdown, and focus_tag columns Adapted dataset: 4,968 rows… See the full description on the dataset page: https://huggingface.co/datasets/Natashadonoh/yoruba-diacritic-restoration-dataset.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes14downloads

Natashadonoh/yoruba-diacritic-restoration-dataset · main · files are served by the source, never re-hosted here