sol-r/historica-corpus
Historica Corpus v2 A monolingual pretraining corpus of 314,438 passages (177M words, 1.1B characters) spanning 15 ancient and historical languages, from 500 BCE to 1750 CE. What's New in v2 2x more passages (314k vs 159k) due to varied-length chunking SGML entity resolution: Middle English þ/ȝ/ð properly rendered (928k entities fixed) PROIEL punctuation: reconstructed from presentation-after attributes TEI apparatus handling: <lem> (main reading) preserved… See the full description on the dataset page: https://huggingface.co/datasets/sol-r/historica-corpus.
This repository belongs to sol-r on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
