Urdatorn/sphragis-olmo1b-adaptation-corpus
Sphragis OLMo-1B adaptation corpus Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek before authorship-language-model training. It contains only OGA whole works whose TLG author occurs in neither Sphragis benchmark. Text has the exact model-facing benchmark surface form: polytonic-aware lowercasing with grc_utils.lower_grc, removal of all editorial punctuation, normalization of whitespace, and removal of consonant-final elision marks. Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.
0109
Publish normalized exclusion-audited OLMo-1B corpus
initial commit
