CoolFace
Datasetpublic

Urdatorn/sphragis-metre

Sphragis Metre Sphragis Metre is the scanned-line companion to Urdatorn/sphragis. It supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line units combining human metrical annotation with uniform automatic dependency annotation. Every curated Hypotactic passage is parsed with the pinned Ericu950/Stoicheia-tagger-parser checkpoint. Tasks There are three task sizes, 1, 5 and 10 lines, on each of three tracks. Track What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.

sourceHugging Faceotherupdated 2d agoView on Hugging Face
0likes336downloads
Dataset Card

Sphragis Metre

Sphragis Metre is the scanned-line companion to `Urdatorn/sphragis`. It supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line units combining human metrical annotation with uniform automatic dependency annotation. Every curated Hypotactic passage is parsed with the pinned `Ericu950/Stoicheia-tagger-parser` checkpoint.

Tasks

There are three task sizes, 1, 5 and 10 lines, on each of three tracks.

TrackWhat its rows arePurpose
verse_*every scanned line of every poetthe main track: all seventeen poets, most data
verse_hexameter_*as verse_*, Hesiod and the two Homeric corpora onlyone metre, one dialect, one period; the same three labels as the verse_hexameter_* track of `Urdatorn/sphragis`
verse_hexameter_all_*every line whose metre is hexameter, from fourteen poetsone metre across eight centuries

The main track mixes lyric, elegiac, iambic and hexameter poetry, so a model can score well on it by telling the metres apart rather than the poets: the metre column alone separates Aeschylus, Pindar and Lycophron from the rest, and it is a Hypotactic annotation, not a fact of the transmitted text. verse_hexameter_all_* removes that shortcut. Every one of its rows carries the metre label hexameter, so Callimachus' elegiac couplets, Theognis' pentameters and Theocritus' Aeolic lyric leave the track together with all of Aeschylus, Pindar and Lycophron, and what is left to tell the poets apart is how each of them writes one metre. verse_hexameter_* narrows that further to the three archaic poets, matching period and dialect as well, and holds exactly the labels and lines of the sentence publication's track of the same name, so the two publications put one question to one set of labels.

The two tracks reuse the main track's lines and the main track's splits, so a model trained on the main track can be scored on them without retraining, and a model trained on them alone can be compared with it.

There is no tragedians_* track here, although the sentence publication has one. Hypotactic scans no Sophocles and no Euripides, so the track would carry the single label Aeschylus.

Configurations

This is the second edition, rebuilt in September 2026 together with `Urdatorn/sphragis`: the 50-line task is gone, every configuration has four splits (train, validation, validation2 and test, the third reserved for choosing among combinations of models), and every split, chunk boundary and seed changed. Scores on the first edition are not comparable with scores on this one.

<!-- tables:begin -->

Configurations

ConfigurationAuthorsRow unitTrain rowsValidation rowsValidation2 rowsTest rows
verse_117line65,3308,6304,2908,630
verse_5175-line chunk13,0661,7268581,726
verse_101710-line chunk6,533863429863
verse_hexameter_13line22,2402,9501,4702,950
verse_hexameter_535-line chunk4,448590294590
verse_hexameter_10310-line chunk2,224295147295
verse_hexameter_all_114line59,3407,8403,8907,860
verse_hexameter_all_5145-line chunk11,8681,5687781,572
verse_hexameter_all_101410-line chunk5,934784389786

Lines per poet, main track (verse_1)

Poet labelTrainValidationValidation2TestTotal
Aeschylus1,5002001002002,000 (2.3%)
Apollonius Rhodius4,3705802905805,820 (6.7%)
Aratus860110501101,130 (1.3%)
Callimachus1,040130601301,360 (1.6%)
Hesiod1,400180901801,850 (2.1%)
Homeric-Iliad11,7701,5607801,56015,670 (18.0%)
Homeric-Odyssey9,0701,2106001,21012,090 (13.9%)
Lycophron1,100140701401,450 (1.7%)
Nicander1,190150801501,570 (1.8%)
Nonnus16,0302,1301,0602,13021,350 (24.6%)
Oppian2,6203501703503,490 (4.0%)
Pindar2,5903401703403,440 (4.0%)
Pseudo-Oppian1,6002101002102,120 (2.4%)
Quintus Smyrnaeus6,6008804408808,800 (10.1%)
Theocritus2,0102601302602,660 (3.1%)
Theognis1,070140701401,420 (1.6%)
Tryphiodorus510603060660 (0.8%)
All labels65,3308,6304,2908,63086,880 (100.0%)

Genre-matched tracks

TrackPoets
verse_hexameter_*Hesiod, Homeric-Iliad, Homeric-Odyssey
verse_hexameter_all_*Apollonius Rhodius, Aratus, Callimachus, Hesiod, Homeric-Iliad, Homeric-Odyssey, Nicander, Nonnus, Oppian, Pseudo-Oppian, Quintus Smyrnaeus, Theocritus, Theognis, Tryphiodorus

<!-- tables:end -->

The main track's configuration prefix is verse. Because every row belongs to this metrical-verse publication, the redundant constant genre column is omitted.

All splits are fixed at the configuration suffix. Within one track, _1, _5 and _10 contain exactly the same source lines in total; seeded remainder removal is determined by the 10-line task and reused across task sizes, and a poet is retained only with at least 100 training, 50 validation, 10 validation2 and 50 test lines of that track. Splits use random seed 917 and are stratified within each poet and work at 75/10/5/10; a line keeps the split it has on the main track, so the tracks never disagree about a line and nothing a model saw in training on one track is test data on another.

python
from datasets import load_dataset

metre = load_dataset("Urdatorn/sphragis-metre", "verse_10")
hexameter = load_dataset("Urdatorn/sphragis-metre", "verse_hexameter_all_10")

The Cynegetica is transmitted under Oppian's name but is not by the poet of the Halieutica. The two works previously shared one label, of which 38% was written by somebody else; Oppian is now the Halieutica alone and the Cynegetica is Pseudo-Oppian.

Pseudo-Oppian is retained rather than excluded, unlike other pseudonymous labels. The Cynegetica is the work of one Syrian poet, demonstrably neither the Cilician poet of the Halieutica nor anyone else published here, so the label denotes a single author and carries no risk of being another label's hand under a second name. Only the poet's name is unknown.

Representation

For convenient human inspection in the Hugging Face dataset viewer, syllables and conllu are the first and second columns in every configuration.

text, conllu, metre and syllables are real nested Arrow columns with one element per constituent line, not JSON packed into strings, so load_dataset returns Python lists and dictionaries with nothing left to parse. Each syllables line object contains its features mapping and its ordered syllables list. Atomic _1 rows therefore contain one line object, while _5 and _10 retain 5 or 10 explicitly distinguishable lines instead of flattening their annotations.

Each conllu unit is a list of sentences, each sentence a list of token objects keyed by the CoNLL-U column names. xpos is not among them: the nine-position tag is folded into feats, which carries all seven Greek tenses as Tense plus Aspect — see "Column structure" in the Sphragis README, whose representation this publication shares. scripts.conllu_units.render_conllu returns canonical CoNLL-U text for a unit.

The line-level features object always contains hiatus, longa, brevia, morae and caesurae. Every syllable is long or short; long syllables contribute 1 mora and short syllables 0.5. longa and brevia count long and short syllables, while hiatus counts syllables carrying that feature.

caesurae holds one object per caesura, giving its position measured two ways: isochronic is the metrical time before it and isosyllabic is the syllable count before it. Earlier releases spread the same information over sparse keys with positional names (isochronic-caesura-3, isosyllabic-caesura-4), which no Arrow type the Hugging Face datasets library understands can express.

json
[
  {
    "features": {
      "hiatus": 1,
      "longa": 2,
      "brevia": 2,
      "morae": 3,
      "caesurae": [
        {"isochronic": 2, "isosyllabic": 3}
      ]
    },
    "syllables": [
      {"text": "μῆ", "quantity": "long", "features": []},
      {"text": "νι", "quantity": "short", "features": []},
      {"text": "ν", "quantity": "short", "features": ["caesura", "hiatus"]},
      {"text": "ἄει", "quantity": "long", "features": []}
    ]
  }
]

Punctuation-delimited contexts are first reconstructed across consecutive metrical lines. All editorial punctuation and symbols are then removed before inference, including marks attached to words (the angle brackets of a supplement such as <ἄλαστ>); a malformed Hypotactic word span that contains internal whitespace is split into separate parser tokens. The cleaned context is parsed in full with Stoicheia and every dependency relation is normalized to the Universal Dependencies v2 inventory (AuxP, for example, becomes case; unknown residual labels fall back to dep), and since the third edition every coordinator is re-attached to the conjunct that follows it: the parser was trained on two treebanks that place cc differently (UD Perseus on the first conjunct, PROIEL on the following one) and did either, and the sentence publication's human layer is reduced to the same rule, so the two publications' trees share one convention. Each sentence tree is then cropped back to its constituent line boundaries. Thus a line break never becomes an artificial parser sentence boundary. If one line contains material from more than one sentence, its conllu unit holds those sentences in order as separate token lists; 2,092 units do. Context-window fallback splits, if required, occur only between lines. Every result is validated with the official CoNLL 2018 loader. Each text line is exactly the forms of its conllu unit joined by single spaces, so a whitespace split of the text aligns one-to-one with the tokens. syllables contains the human syllable transcription and quantities plus two uniformly derived word-boundary features, caesura and hiatus, which are the only syllable features published. The other upstream Hypotactic syllable classes are per-file annotator conventions rather than metrical facts (lbn, lbp and correption occur only in the Homer, Quintus and Nonnus files, fixme and iambwrong only in Aeschylus, and context-menu-active is web-interface state), so passing them through would have named the source file on every syllable; they are dropped. caesura is added to the final syllable of every non-line-final Hypotactic span.word, independently of incidental HTML formatting. hiatus is likewise derived uniformly: upstream hiatus classes are removed, and the feature is added to the first word's final syllable exactly when grc_utils.vowel returns true for both that word's final character and the next word's initial character. Line-final syllables never receive either derived boundary feature. The redundant scansion field is not published.

Source syllable transcriptions are preserved rather than silently corrected. A line whose syllable transcription does not spell its text (upstream transcriptions occasionally carry a word the line does not, or cut one short) or that contains a syllable of unknown quantity is excluded instead; the counts are recorded under hypotactic_lines_excluded_by_reason in metadata/build_report.json, and every published line's syllables spell its text exactly.

text_distorted is the content-masked text of every line: the same tokens with every one the parser tags NOUN, PROPN, ADJ or VERB replaced by that tag, the sentence publication's style-only view carried over (see "The content-masked text" in the Sphragis README).

The scheme-only floors of this publication are in `metadata/leakage_report_conllu.json`: logistic regressions over the rates of UPOS, DEPREL and FEATS values, head direction, binned head distance and function words heading anything, which see no Greek at all, on every track. There is one annotation source here, so the source probe is trivial and is recorded as such; the author probe is the floor a model must clear.

syntax_annotation is uniformly predicted, and treebank_source is uniformly stoicheia_tagger_parser, in atomic rows and chunks. source_records pins Stoicheia revision cd8ae1658c364874c3b6f4df37bd1d46313e3cea. The prior gold trees are not published as model input. Where exact alignment exists, they are used only as a diagnostic reference for token-boundary, lemma, UPOS, dependency-relation, UAS, and LAS measurements recorded in metadata/build_report.json.

Sphragis Metre publishes the same model-facing representation as `Urdatorn/sphragis`, so one model can read both without changing its feature extraction. CoNLL-U documents contain token rows only; the text column carries the text. MISC is empty, and any head repair is recorded per token in native_syntax. Word forms, lemmas and text are lowercased with grc_utils.lower_grc, elision is written the same way everywhere, and UPOS, XPOS, FEATS and DEPREL use the shared inventories documented in the Sphragis README. Full audit provenance remains available in dedicated metadata columns.

This publication is annotated by one parser throughout, so it does not carry the cross-treebank confound that motivates the shared inventories; they are applied here for comparability, not for leakage control.

The source revisions, split manifest, exclusions, and validation results are recorded under metadata/.