CoolFace
Datasetpublic

TigreGotico/arabic-dialects-gold20

arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes699downloads
24 commits on main
b07d3212mo ago

Improve dataset card: intended use, limits, noise/quality profile

Jarbas
54a707a2mo ago

arbitrate the 14 remaining flagged blind-verification rows (C9, cited); card aggregates updated

Jarbas
11f2cd72mo ago

Refresh dataset card: PER tables + Rijal Almaʿ qaf correction note

Jarbas
9473eb22mo ago

Refresh card: PER tables + Rijal Almaʿ qaf->[ɡ] correction note

Jarbas
603e2ab2mo ago

Saudi wave: Sharqiyya jim->[dʒ], Najd affricate revert, Rijal Almaʿ qaf->[ɡ] (o2i #567); dagger-alif rooting; regen ipa_o2i

Jarbas
11c38d92mo ago

card: drop validation-pack reference

Jarbas
32113092mo ago

remove review sheets from the data repo (loader pollution); relocated to TigreGotico/dialect-validation-packs

Jarbas
6b37d022mo ago

sync to current engine drafts + PER tables

Jarbas
a650bb72mo ago

engine PER tables: vocalized o2i 0.014 / arbtok 0.024 / espeak 0.232; bare o2i 0.291 / arbtok 0.189 / espeak 0.357

Jarbas
c8ec53a2mo ago

card: bare-input PER table (o2i 0.291 / arbtok 0.217 / espeak 0.357) — the deployment-relevant comparison

Jarbas
c2413a22mo ago

native-review pack: per-lect sheets with non-linguist romanization and two-question protocol

Jarbas
62f365f2mo ago

timeless self-contained dataset card

Jarbas
b94deec2mo ago

card: three-engine PER table (o2i 0.014 / arbtok 0.055 / espeak-ar 0.232) with honest interpretation

Jarbas
af9a8342mo ago

verification COMPLETE: 33/33 lects, 654 rows scored, 127 disputes arbitrated, mean agreement 0.837; final methodology card

Jarbas
1f6259c2mo ago

card: blind-verification methodology section

Jarbas
083c4b32mo ago

verification round: 8 lects blind-verified (2 judges each) + arbitrated — EG MA SY najd JO KW IQ LB; verification/judge_agreement columns; ~40 C9 consensus corrections

Jarbas
e7b634a2mo ago

v2.1: ween deltas resolved at source (#459/#460/#462) — engine and gold now agree on the five where-rows

Jarbas
8c462422mo ago

Fable-gold v2: full row-by-row review pass — 322/660 rows corrected across 8 classes (C1-C8); SD register fix at source

Jarbas
e0be73f2mo ago

Fable-gold v1: ipa = corrected gold, ipa_o2i = engine draft; 295 rows corrected (C1 Allah-forms, C2 Gulf ween, C3 clitic destressing, C4 mater artifacts), each cited in fable_corrections

Jarbas
85653472mo ago

dataset card: full methodology, provenance, and limitations

Jarbas
b69964f2mo ago

gold-20 COMPLETE: 660 rows, 33 lects x 20 Fable-verified sentences

Jarbas
2a5e06e2mo ago

gold-20 rollout: 33 lects incl. 5 grouping nodes and 3 new Saudi leaves; 646 rows, 32 lects at 20

Jarbas
1566c672mo ago

arabic_tts gold: 28 lects, imperative round + gold-20 rollout (ar, arb, ar-EG at 20)

Jarbas
96e2cc12mo ago

initial commit

Jarbas