Natashadonoh/yoruba-diacritic-restoration-dataset
Yorùbá Diacritic Restoration Dataset Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text. Dataset Details Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded) Raw dataset: 8,365 rows, annotated with tone_pattern, harmony_class/harmony_breakdown, and focus_tag columns Adapted dataset: 4,968 rows… See the full description on the dataset page: https://huggingface.co/datasets/Natashadonoh/yoruba-diacritic-restoration-dataset.
task_categories:
- text2text-generation ---
Yorùbá Diacritic Restoration Dataset
Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text.
Dataset Details
- Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded)
- Raw dataset: 8,365 rows, annotated with
tone_pattern,harmony_class/harmony_breakdown, andfocus_tagcolumns - Adapted dataset: 4,968 rows, generated via Adaption Labs' adaptation pipeline (Instruction Tuning recipe)
Columns
Task
Given yoruba_stripped as input, restore the correctly diacritized form (yoruba) — a standard diacritic-restoration task in Yorùbá NLP.
License
CC BY-NC 4.0, inherited from the MENYO-20k source corpus.
Author
Areo Mafadesere Natasha
