CoolFace
Datasetpublic

Urdatorn/sphragis-olmo1b-adaptation-corpus

Sphragis OLMo-1B adaptation corpus Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek before authorship-language-model training. It contains only OGA whole works whose TLG author occurs in neither Sphragis benchmark. Text has the exact model-facing benchmark surface form: polytonic-aware lowercasing with grc_utils.lower_grc, removal of all editorial punctuation, normalization of whitespace, and removal of consonant-final elision marks. Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes109downloads
2 commits on main
b71f2c51mo ago

Publish normalized exclusion-audited OLMo-1B corpus

Urdatorn
20d7af51mo ago

initial commit

Urdatorn