CoolFace
Datasetpublic

Salesteq/arabic-dialects-gold20

arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes790downloads
Dataset Card

arabic-dialects-gold20

660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row.

Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification, judge_agreement

  • `ipa` — the gold column: reviewed, corrected IPA of the sentence as a speaker of the lect says it.
  • `ipa_o2i` — the draft the gold was corrected from: the output of orthography2ipa, whose specs encode each lect's cited rules. Where the two columns differ, the delta is a documented engine correction; where they agree, the engine's cited rules already produce the gold value.
  • `fable_corrections` — every correction applied to the draft, each with its citation.
  • `verification` / `judge_agreement` — per-row blind-verification outcome (see below).

Lects: MSA (ar), Classical (arb), 26 dialect leaves (Egyptian, Sudanese, Chadian, Nigerian/Shuwa, 4 Levantine, Iraqi gilit + Muslawi qeltu, 5 Gulf, Omani, Yemeni/Ṣanʿānī, 5 Saudi incl. Qassimi/Rijāl Almaʿ/Sharqiyya, 5 Maghrebi incl. Ḥassāniyya), and 5 dialect-group nodes (ar-x-{gulf,levantine,maghrebi,mashriqi,peninsular}) carrying pan-group register.

What this dataset is — and is not

A synthetic, LLM-authored, literature-grounded gold set. No native speaker has validated it. It approximates a per-dialect, sentence-level Arabic IPA corpus at a coverage no native-validated resource offers, with every construction decision documented and every correction cited. It suits TTS-frontend validation, G2P regression testing, and engine comparison; it is not ground truth for dialectology claims.

Noise / quality profile

Text-only — no audio in this repository, so no bandwidth/DNSMOS/UTMOS profile applies. Quality here means transcription/phonemization reliability, tracked per row via verification/judge_agreement and per-correction via fable_corrections.

How each row is made

  1. 1.Authoring. Sentences are authored by Claude (Anthropic; a lead model with per-dialect workers) on 14 shared semantic frames (weather+negation, price+number, appointment time, family/residence, directions imperative, food evaluation, where-question, how-many, conditional, past shopping narrative, future travel, phone, greeting pair, proverb/closing) plus 6 per-lect coverage sentences. Shared frames keep lects comparable; every content word must be attested for the lect in its cited literature — MSA-with-dialect-reflexes counts as a defect. Sentences are sandhi-rich by design (article liaison, waṣl chains, tāʾ-marbūṭa liaison, cross-word assimilation, pausal endings).
  2. 2.Vocalization. Dialect tashkeel is hand-set following the set's written conventions (sun-letter assimilation as shadda on the sun letter, written waṣl/article contractions, pausal final forms).
  3. 3.Drafting. ipa_o2i is generated by orthography2ipa from the vocalized sentence.
  4. 4.Correction. The gold ipa is the draft corrected against the lect's primary literature. Corrections are classed and cited in fable_corrections (classes include stress placement on clitics and fused proclitics, vowel quantity the orthography under-determines — demonstratives, the Allah lam — reduction vocalism, article forms, loan-class exemptions, construct-state /t/, and segment repairs). When a correction reveals an engine defect rather than a transcription slip, it is also filed against the engine, so the ipa/ipa_o2i delta doubles as an engine bug ledger. A late correction of note: Rijāl Almaʿ qāf is [ɡ], not the [q] earlier assumed (Alfaifi 2024 pp.177-178, quoting Asiri's field data); the genuine Tihama split for this lect is the ẓāʾ reflex ([ɬˤ] vs [ðˤ]), not qāf.
  5. 5.Feature tags are machine-verified predicates (qaf/jim/interdental/ kaf/hamza/tamarbuta/sunassim on the orthography; emphatic/pharyngeal/ geminate/long_vowel/diphthong on the IPA) plus shape tags (statement/question/negation/imperative/number). Citation ids in notes are machine-checked against each lect's spec bibliography.

Blind verification

Every lect's sentences are independently re-transcribed by two blind judges — separate LLM sessions that receive only the vocalized sentences and the lect name, may not read the gold or call the engine, and derive their IPA from the dialectological literature, each anchored on a different primary source. A scoring harness normalizes (stress, lax/tense notation) and computes segment-level agreement. Buckets:

  • unanimous — both judges agree with the gold;
  • majority — one does;
  • disputed-arbitrated-judges — both judges diverged together and arbitration against the primary sources corrected the gold (the citation is in fable_corrections);
  • disputed-arbitrated-gold — the judges were wrong and the note says why;
  • disputed / noisy — flagged, not yet arbitrated.

Aggregate over the scored rows: 390 unanimous, 123 majority, 141 arbitrated disputes (85 corrected toward the judges, 56 resolved for the gold), none left flagged; mean judge agreement 0.837 (per-lect means 0.78–0.89). Group-node files verify weakly by construction: a pan-group target is under-determined, and judges legitimately pick different leaf mainlines. Rows edited after scoring carry no verification value and say so.

Engine PER against this gold (full dataset, no caps)

Phoneme error rate of three systems against the gold ipa column (segment-level edit distance, stress stripped, lax/tense folded).

Vocalized input (the sentence column — full tashkeel)

lecto2iarbtokespeak (ar)
ar-AE0.020.0170.279
ar-BH0.0120.0120.282
ar-DZ0.0180.0180.192
ar-EG0.0020.0030.27
ar-IQ-x-qeltu0.0020.0020.199
ar-IQ0.0050.0050.295
ar-JO0.00.00.191
ar-KW0.00.00.342
ar-LB0.0220.0180.278
ar-LY0.0060.0060.228
ar-MA0.0510.0510.238
ar-MR0.0050.0050.244
ar-NG0.0090.0090.209
ar-OM0.0030.0030.192
ar-PS0.020.0210.21
ar-QA0.00.00.3
ar-SA-x-hejaz0.0130.0140.213
ar-SA-x-najd0.00.00.216
ar-SA-x-qassim0.0070.0070.249
ar-SA-x-rijal-alma0.0040.0070.203
ar-SA-x-sharqiyya0.0110.0110.205
ar-SD0.0210.0210.213
ar-SY0.0550.0550.289
ar-TD0.0310.0310.212
ar-TN0.0010.0030.247
ar-YE0.00.00.163
ar-x-gulf0.0020.0020.287
ar-x-levantine0.0020.0020.266
ar-x-maghrebi0.00.00.18
ar-x-mashriqi0.0080.0080.2
ar-x-peninsular0.0070.0060.177
ar0.00.0570.16
arb0.00.0560.157
mean0.0100.0140.230

Bare input (the raw column — what a TTS actually receives)

lecto2iarbtokespeak (ar)
ar-AE0.3120.2040.398
ar-BH0.2540.1870.403
ar-DZ0.2130.2120.324
ar-EG0.2930.1530.384
ar-IQ-x-qeltu0.2720.160.317
ar-IQ0.2750.1960.38
ar-JO0.2880.1820.352
ar-KW0.2820.2010.408
ar-LB0.3380.2130.375
ar-LY0.2690.1990.385
ar-MA0.2380.2370.342
ar-MR0.2830.2050.387
ar-NG0.3080.1810.357
ar-OM0.2780.1340.301
ar-PS0.2960.1710.354
ar-QA0.2790.1750.401
ar-SA-x-hejaz0.2830.1380.325
ar-SA-x-najd0.2970.1520.334
ar-SA-x-qassim0.2740.1490.395
ar-SA-x-rijal-alma0.2860.1350.325
ar-SA-x-sharqiyya0.3170.1520.321
ar-SD0.2910.1590.35
ar-SY0.3140.2260.372
ar-TD0.2970.1660.35
ar-TN0.2420.1730.36
ar-YE0.2540.0910.285
ar-x-gulf0.290.1850.416
ar-x-levantine0.2980.2070.389
ar-x-maghrebi0.2340.2040.325
ar-x-mashriqi0.2910.1660.294
ar-x-peninsular0.2670.1360.282
ar0.390.2110.382
arb0.430.1830.357
mean0.2890.1770.355

How to read this honestly:

  • o2i shares its cited rules with the gold's draft layer, so its vocalized PER is not independent accuracy — it is the size of the correction layer for that lect.
  • arbtok runs its full TTS pipeline. On vocalized input its distance beyond o2i is dominated by the pausal-register choice on ar/arb (whose gold keeps full iʿrāb style); on the dialect lects the two columns are close. On bare input its diacritization (dialect-aware rawi-lattice fusion) and stem lexicon restore what raw o2i must guess.
  • espeak-ng uses its single ar (MSA) voice for every lect — the column measures how far MSA speech is from each dialect, not an espeak defect; the gradient behaves as dialectology predicts.
  • The bare-input table is the deployment-relevant comparison.

Sources grounding the registers (per-lect authorities)

Badawi & Hinds 1986, Mitchell 1956 (Egyptian); Cowell 1964, Almbark & Hellmuth 2015 (Levantine); Al-Wer 2020, Fadda 2016 (Ammani); Cotter 2016 (Gazan); Erwin 1963, Blanc 1964, Jasim 2020 (Iraqi); Holes 2004, Al-Balushi 2016, Alshammari 2026, Johnstone 1967 (Gulf/Omani); Watson 2002 (Ṣanʿānī and pan-Arabic phonology); Ingham 1994 (Najdi); Alhoody 2019, Al-Rojaie 2013 (Qassimi); Omar 1975, Abdoh 2010 (Hejazi); Watson & Al-Azraqi 2011, Asiri 2009 (Rijāl Almaʿ); Al-Taisan 2022 (Eastern Province); Dickins 2007 (Sudanese); Owens 1993/2006, Kaye 1976, Jullien de Pommerol 1999 (Chadian/Nigerian); Harrell 1962, Heath 2020 (Moroccan); Singer 1984, Gibson 2009, La Rosa 2021 (Tunisian); Boucherit 2002, Marçais 1902 (Algerian); Pereira 2010/2011, Benkato 2020 (Libyan); Taine-Cheikh 2007, Cohen 1963 (Ḥassāniyya).

Known limitations

  • Not native-validated. The verification layer is a second and third independent LLM derivation with arbitration — the strongest check available without native speakers, and not a substitute for one.
  • The IPA reflects each lect's mainline register as its spec models it; documented in-lect variation (Baḥārna vs Sunni Bahraini, Bedouin vs urban Chadian, madani vs koine Ammani qāf) collapses to one value, stated in the notes.
  • Where a sentence would require a construction the engine mishandles, the set uses an equivalent construction instead and the engine gap is recorded in the notes — coverage of those seams lives in the correction ledger, not in silent workarounds.
  • The engine-PER table above is partly engine-similarity for o2i/arbtok (shared cited rules with the draft layer); only the espeak column is a fully independent system.

Source of truth: orthography2ipa/data/gold/arabic_tts/ in the o2i repo, which pins ipa_o2i in CI. Sibling set: TigreGotico/gold20-code-switch.