Abdullah-afify/egyptian-names
Egyptian Names & Onomastic Intelligence Dataset From 15.88M+ Raw National Records to an Empirical Onomastic and Linguistic Engine This repository hosts the complete, multi-phase statistical and linguistic dataset powering egy-names, the production onomastic intelligence engine for contemporary Egyptian naming traditions. Dataset Pipeline Overview Egyptian names follow an unbroken patronymic lineage chain ($Personal + Father + Grandfather + Ancestor… See the full description on the dataset page: https://huggingface.co/datasets/Abdullah-afify/egyptian-names.
Link the egy-names-fallback-classifier model
Sync 0.3.5 catalog fixes: collision resolution, compound-spacing corrections, add is_personal_name/is_low_confidence audit columns
Sync 0.3.5 catalog fixes: collision resolution, compound-spacing corrections, add is_personal_name/is_low_confidence audit columns
Update README.md (sanitized, 14D schema, zero emojis)
Update dataset documentation and schema for v0.3.1 (14D onomastic intelligence)
Update v0.3.0 with authentic bilingual Dalla' (tashkeel+ipa) and Egyptian public figures
Update v0.3.0 with authentic bilingual Dalla' (tashkeel+ipa) and Egyptian public figures
Release v0.3.0 with 11-Dimensional Onomastic & Phonetic Enrichment (44.6K names)
Release v0.3.0 with 11-Dimensional Onomastic & Phonetic Enrichment (44.6K names)
docs: Update Hugging Face dataset README with complete 44.6K lexicon schema & metrics
v0.2.1: 100% complete Gemini-validated dataset (data/slot_distributions.csv)
v0.2.1: 100% complete Gemini-validated dataset (data/slot_distributions.parquet)
v0.2.1: 100% complete Gemini-validated dataset (data/final_canonical_names.csv)
v0.2.1: 100% complete Gemini-validated dataset (data/final_canonical_names.parquet)
v0.2.1: 100% complete Gemini-validated dataset (data/names.csv)
v0.2.1: 100% complete Gemini-validated dataset (data/names.parquet)
v0.2.1: 100% complete Gemini-validated dataset (README.md)
v0.2.1: update data/corrections.csv — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/corrections.csv — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/corrections.csv — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/slot_distributions.csv — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/slot_distributions.parquet — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/final_canonical_names.csv — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/final_canonical_names.parquet — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/names.csv — 44,626 names, Phase 8 LLM annotations
v0.2.1: update data/names.parquet — 44,626 names, Phase 8 LLM annotations
v0.2.1 final: data/slot_distributions.parquet — 44,626 names, 100% slot coverage
v0.2.1 final: data/final_canonical_names.parquet — 44,626 names, 100% slot coverage
v0.2.1 final: data/names.parquet — 44,626 names, 100% slot coverage
v0.2.1: Update data/corrections.csv — 40,169 names, 23,457 corrections
v0.2.1: Update data/corrections.parquet — 40,169 names, 23,457 corrections
v0.2.1: Update data/slot_distributions.parquet — 40,169 names, 23,457 corrections
v0.2.1: Update data/final_canonical_names.parquet — 40,169 names, 23,457 corrections
v0.2.1: Update data/names.parquet — 40,169 names, 23,457 corrections
docs: update Afify Corporation enterprise scope (software, hardware, media, advanced tech & AI)
Add data/phase0_raw_full_names.parquet
Add data/slot_distributions.csv
Add data/phase1_segmented_chains.parquet
Add data/phase1_segmented_chains.parquet
Add data/phase1_segmented_chains.parquet
Add data/phase0_raw_full_names.parquet
Add data/slot_distributions.csv
Add data/slot_distributions.parquet
Add data/phase3_spelling_corrections.csv
Add data/phase2_token_frequencies.csv
Add data/phase3_spelling_corrections.parquet
Add data/phase2_token_frequencies.csv
Add data/phase2_token_frequencies.parquet
Add data/names.parquet
Add data/names.parquet
