CoolFace
Datasetpublic

TigreGotico/wikipron-restored-orthography

WikiPron with restored orthography A pinned mirror of every pronunciation scrape in CUNY-CL/wikipron, with a second orthography column holding the headword English Wiktionary actually displays. The defect WikiPron pairs a pronunciation with the English Wiktionary MediaWiki page title. For a number of languages the title is not the word the page displays, because the language's style policy keeps diacritics out of titles and puts them back only on the headword… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/wikipron-restored-orthography.

sourceHugging Facecc-by-sa-4.0updated 26d agoView on Hugging Face
0likes429downloads
Dataset Card

WikiPron with restored orthography

A pinned mirror of every pronunciation scrape in CUNY-CL/wikipron, with a second orthography column holding the headword English Wiktionary actually displays.

The defect

WikiPron pairs a pronunciation with the English Wiktionary MediaWiki page title. For a number of languages the title is not the word the page displays, because the language's style policy keeps diacritics out of titles and puts them back only on the headword line. `Wiktionary:About Middle High German` states it plainly:

certain letters with diacritics (ë ā ē ī ō ū ȥ) are not used in article titles, but are used when displaying the word
this does not apply to the umlauted vowels ä, ö, ü, æ, œ, which are treated as separate letters and thus appear in titles like any other letter

The pronunciation still transcribes what the title threw away. A grapheme-to-phoneme system scored on those rows is asked to recover tone, vowel length or vocalisation from an input that no longer writes them, and the resulting error rate measures the input, not the system.

The recovery

restored_orthography holds the headword Wiktionary renders for that language, read from the MediaWiki API. Nothing is guessed:

  • —A page holds one section per language. Adam renders an English Adam and an Ewe Ádàm. The wikitext is cut to the target ==Language== section before any headword template is read, and the rendered headword is matched on its lang attribute.
  • —A page with no headword override, no section for the language, or two headwords that disagree contributes an empty cell. One unpointed title often spells several distinct words — aba renders abà, abá, àba and àbá — and the row is refused rather than assigned a winner.
  • —A restored form must differ from the title only in diacritics. If the base-letter spine changes, the cell stays empty.

An empty restored_orthography therefore means one of: the language was screened and is not affected, the language is affected but has not been restored yet, or the row was refused. The manifest says which.

Schema

orthography, restored_orthography, ipa, tab separated, one header row. Columns one and three are WikiPron's own two columns, byte for byte. File names are WikiPron's: <iso639-3>_<script>_<broad|narrow>, with a dialect segment where upstream carries one.

Consumers choose the column. Score orthography to reproduce a WikiPron number; score restored_orthography on the rows where it is filled to measure a system against an input that encodes the contrast.

Screen

Each language was screened once against Wiktionary:About <Language>, then checked against the data. A policy saying macrons never reach titles predicts zero macrons in the scraped orthography, and the count either bears that out or it does not.

Latin is why the second half is not optional. Its policy says the page name "should not contain diacritical marks", in almost the same words as Old English's, and yet the scraped Latin words carry 1.31 macrons per row against Old English's 0.00005. The disagreement was noticed by reading the two together, and Latin is recorded as not_affected on that basis. The build also enforces it: a confirmed verdict whose data carries the marks above a stray rate is downgraded automatically, so the next language like Latin does not depend on someone noticing.

A verdict is a snapshot against a live wiki, and the wiki moves. Mandarin is recorded no_policy_page, yet Wiktionary:About Mandarin redirects to Wiktionary:Chinese entry guidelines, which exists. That page describes no title/headword split, so the verdict stands, and the mismatch is what drift looks like: expect some no_policy_page to have grown a page since. This is the same reason the scrape itself is pinned — only the snapshot is stable.

Negative results are recorded on purpose. They are the reason nobody has to screen these languages again.

Snapshot

CUNY-CL/wikipron commit d282e848a211ea31cfd730f0ced8bc8cdab9e83d, dated 2026-07-23, path data/scrape/tsv. 543 files.

WikiPron publishes its scrapes on a branch it keeps editing, so a benchmark reading it directly cannot tell a system change from an upstream edit. Every file here is copied from that one commit and manifest.json records each file's upstream SHA-256, which is what makes a refresh a diff rather than a guess.

Screen verdicts and coverage

A confirmed language with no restored rows has not been run. “not attempted” in the table and restoration_attempted: false in manifest.json say so explicitly, because a language that was run and recovered nothing would mean a broken lookup, not an empty Wiktionary, and the two must not look alike. The build refuses to publish the second case.

LanguageRowsVerdictRestoredPolicy
ajp6543confirmednot attempted“Page titles never include diacritics, including shadde.”
ang85117confirmednot attempted“Consequently, Old English entries here should be without diacritical marks in the page title.”
apc805confirmed103“Page titles never include diacritics, including shadde.”
car447confirmed11“Irregular stresses should not be marked in entry names, but should be marked with an acute accent in alternative display parameters.”
ceb8131confirmednot attempted“Diacritics are normally not used in written Cebuano, but are used for headwords in most Cebuano dictionaries to distinguish homographs.”
dum222confirmed49“Diacritics should not be used in entry names.”
evn153confirmed104“Long vowels are not represented in the entry name, but should always be indicated in the headword with a macron.”
gmh1724confirmed408“Certain letters with diacritics (ë ā ē ī ō ū ȥ) are not used in article titles, but are used when displaying the word.”
gml175confirmed83“When creating a Middle Low German entry, the head (but not the actual page title) should follow the tradition of Middle Low German research to mark originally short vowels with a macron and original long vowels and diphthongs with a circumflex.”
goh199confirmed61“This macron is to be used only for display, not in entry names.”
grc198102confirmednot attempted“Entry names do not have macrons or breves.”
hau4121confirmednot attempted“Diacritical marks should not be used in page titles, but should always be used in headwords.”
hbs102732confirmednot attempted“In the headword line, such accent marks should be specified as alternative displays, by means of the head= parameter.”
heb8088confirmed4307“Do not use niqqud (vowel points) in page names, but do include it in headword-line templates.”
nci2373confirmednot attempted“Long vowels are marked with macrons only within the text of pages, not in page names.”
nya1624confirmed675“The circumflexed letter ŵ should not be in entry titles, but it should be in headword lines. Tones should also be marked in headword lines, using the acute to mark high tones.”
okm808confirmed414“The entry titles for Middle Korean terms should be written in the Hangul script as invented by Sejong, without tone marks.”
osx273confirmed101“This macron is to be used only for display, not in entry names, so the additional parameter that is available in many templates should be used to change the displayed form without affecting the link.”
pam1856confirmed885“Headwords should have diacritics as a pronunciation guide.”
pan5104confirmednot attempted“Similarly, diacritics can be utilised in the page body, but should also be avoided in page titles.”
sga7229confirmed120“This parallels Wiktionary's approach to Latin and Old English, where macrons are used in display but not in page titles.”
tgl61357confirmednot attempted“Diacritics are normally not used in written Tagalog, but are used for headwords in most Tagalog dictionaries to distinguish homographs.”
yor4937confirmed3257“The underdot vowels, ẹ and ọ, should be used in page titles, but the tones should be marked in the headword line.”
yrk455confirmed121“Long and short vowel diacritics are to be supplied in the headwords of the appropriate entries.”
acm108confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
acw3446confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
afb763confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
ara17678confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
ary2168confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
arz1380confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
ayl166confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
ewe479confirmed, no policy page430no About page; the scraped words are all but untoned and the transcriptions mark tone throughout
msa14359confirmed, no policy pagenot attemptedno statement on the About page; not one short-vowel mark appears in any scraped word, while the transcriptions are fully vocalised
amh478not affectednot attemptedno title/display split; the vowel is written into the syllabary character itself, so nothing is stripped
cat98885not affectednot attempted“It is encouraged to use sort=(the page name without diacritics) in headword-line templates.”
cop881not affectednot attempted“Other sporadically appearing diacritics such as circumflexes, acute accents, and the like should generally not be used in entry names or headword lines of main lemmas.”
enm18272not affectednot attempted“Middle English entries here should be without diacritical marks, whether in the page title or within the entry itself.”
fas50495not affectednot attemptedno title/display split; short vowels are absent from the displayed headword too, so there is nothing to restore
got2230not affectednot attempted“As in other old languages, macrons are not used in these entry names, although the got-rom template allows a head= parameter to display them if necessary.”
haw4127not affectednot attempted“Macrons and the okina should always be used in page titles.”
lat87895not affectednot attempted“For these reasons, the page name for Latin entries should not contain diacritical marks.”
nld118103not affectednot attempted“On Wiktionary, entry names containing stress marks are permitted only where they are used to distinguish one word from another.”
ota385not affectednot attemptedno title/display split; short vowels are absent from the displayed headword too, so there is nothing to restore
vec141not affectednot attempted“The headword should always match the entry title — no additional diacritics should be added.”
yid5895not affectednot attempted“Normal entries should have titles that use all appropriate diacritical marks, including subscript vowels and other niqqudim.”

Of the remaining languages, 164 have no Wiktionary:About <Language> page at all and 139 have one that says nothing about diacritics in titles. Absence of a policy is an answer too: there is no documented title/display split to recover, and their restored_orthography column is empty. manifest.json carries a verdict for every language, with the marks each policy names and how often they occur in the scraped words.

Limitations

Coverage is partial and always will be. Restoration costs one page render per word, so the largest affected languages are screened and recorded but not restored; their restored_orthography column is empty throughout, and the manifest marks them.

This is English Wiktionary only. It inherits WikiPron's own quality issues without correcting any of them — crowd-sourced transcriptions, uneven transcription traditions inside one language, and entries whose headword line and pronunciation line were written independently.

Two results are easy to over-read:

  • —Ewe restores to a near-zero error rate against a grapheme-to-phoneme system. Tone-marked Ewe spelling is very nearly a transliteration of its own phonemic transcription, so the row is easy rather than the system good.
  • —Middle High German gets worse after restoration. The restored input carries ⟨ë ā ē ī ō ū ȥ⟩, which the system tested had no rules for at all. A lossy input had been hiding a real gap: while the input never contained those letters, nothing could reveal they were unmapped.

Licence and attribution

The text is English Wiktionary content, licensed CC BY-SA 4.0, and this dataset carries the same licence. WikiPron's Apache 2.0 covers its scraper, not the scraped text. Attribute both English Wiktionary and CUNY-CL/wikipron.