Urdatorn/sphragis-olmo1b-adaptation-corpus
Sphragis OLMo-1B adaptation corpus Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek before authorship-language-model training. It contains only OGA whole works whose TLG author occurs in neither Sphragis benchmark. Text has the exact model-facing benchmark surface form: polytonic-aware lowercasing with grc_utils.lower_grc, removal of all editorial punctuation, normalization of whitespace, and removal of consonant-final elision marks. Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.
Sphragis OLMo-1B adaptation corpus
Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek before authorship-language-model training. It contains only OGA whole works whose TLG author occurs in neither Sphragis benchmark.
Text has the exact model-facing benchmark surface form: polytonic-aware lowercasing with grc_utils.lower_grc, removal of all editorial punctuation, normalization of whitespace, and removal of consonant-final elision marks. Splits are made over whole works before tokenization. Parquet files are the human-inspectable normalized corpus; NumPy arrays are the exact packed token blocks used by training.
Splits
Excluded benchmark authors
All revisions, fingerprints, record IDs, normalization counts, and the comparison with the preceding exclusion list are in manifest.json.
Author
Albin Thörn Cleland, Lund University. ORCID 0009-0003-3731-4038.
