CoolFace
Datasetpublic

ebrinz/hyper-glyphy-artifacts

hyper-glyphy — trained artifacts mirror Companion artifact store for github.com/ebrinz/hyper-glyphy: cross-lingual word-embedding alignment for six ancient languages (Sumerian, Akkadian, Hittite, Ancient Greek, Egyptian, Sanskrit) into GloVe 300d and whitened-EmbeddingGemma 768d English spaces. Everything here is computed output of the pipelines in the GitHub repo, mirrored so results can be reproduced exactly without retraining (FastText training is non-deterministic, so… See the full description on the dataset page: https://huggingface.co/datasets/ebrinz/hyper-glyphy-artifacts.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes95downloads
Dataset Card

hyper-glyphy — trained artifacts mirror

Companion artifact store for github.com/ebrinz/hyper-glyphy: cross-lingual word-embedding alignment for six ancient languages (Sumerian, Akkadian, Hittite, Ancient Greek, Egyptian, Sanskrit) into GloVe 300d and whitened-EmbeddingGemma 768d English spaces.

Everything here is computed output of the pipelines in the GitHub repo, mirrored so results can be reproduced exactly without retraining (FastText training is non-deterministic, so retrained vectors differ slightly from the published numbers). Raw third-party corpora are NOT included — each slot's README in the GitHub repo documents the fetch steps (DCS, Diorisis, ORACC, ETCSL, TLHdig, Cologne CDSL, etc.).

Layout

Mirrors the GitHub repo's gitignored paths:

shared/models/                      English caches: EmbeddingGemma 768d (gloss/bare,
                                    raw + whitened) and whitening transforms
languages/<slot>/models/            FastText model + .vec, fused 1536d npz,
                                    Ridge weights, Procrustes maps
languages/<slot>/data/processed/    merged/cleaned corpora (per-text doc IDs),
                                    english_anchors.json, anchor stats
languages/<slot>/data/dictionaries/ derived gloss lexica (LSJ, Monier-Williams)
languages/<slot>/results/           alignment results + eval-suite artifacts
languages/<slot>/final_output/      production aligned vectors + vocab

Note: in akkadian/greek/hittite/sumerian, FastText files carry the legacy name fasttext_sumerian.* (cloned script kept the output filename); Sanskrit and Egyptian use their own slot names.

Reproducing

bash
git clone https://github.com/ebrinz/hyper-glyphy
cd hyper-glyphy
hf download ebrinz/hyper-glyphy-artifacts --repo-type dataset --local-dir .
# then e.g.:
python languages/sanskrit/scripts/09b_align_gemma.py --mode whitened
python shared/scripts/procrustes_align.py --slot sanskrit

The English GloVe cache (glove.6B.300d.txt, Stanford NLP, PDDL) ships under languages/sumerian/data/processed/; other slots reference it — recreate the symlinks or copy it (see each slot README).

Licensing (mixed — per source)

  • Trained vectors, weights, maps, results (our computation): CC BY 4.0.
  • Sanskrit-derived files (languages/sanskrit/data/dictionaries/mw_glosses.json, english_anchors.json and downstream anchors): derived from the Cologne CDSL digitization of Monier-Williams (1899), CC BY-NC-SA 3.0 — non-commercial, share-alike, attribution: The Sanskrit Library / Thomas Malten / Universität zu Köln. Sanskrit corpora derive from the Digital Corpus of Sanskrit (Oliver Hellwig), CC BY 4.0.
  • Greek-derived gloss files: from Perseus LSJ (CC BY-SA 4.0, Perseus Digital Library).
  • Other slots' processed corpora: derived from ORACC / ETCSL / TLHdig / Diorisis / TLA — see the GitHub repo's per-slot READMEs for source attributions.
  • glove.6B.300d.txt: Stanford GloVe, Public Domain Dedication and License v1.0.

If you use these artifacts, cite the underlying sources per slot (see GitHub READMEs).