CoolFace
Datasetpublic

sol-r/historica-corpus

Historica Corpus v2 A monolingual pretraining corpus of 314,438 passages (177M words, 1.1B characters) spanning 15 ancient and historical languages, from 500 BCE to 1750 CE. What's New in v2 2x more passages (314k vs 159k) due to varied-length chunking SGML entity resolution: Middle English þ/ȝ/ð properly rendered (928k entities fixed) PROIEL punctuation: reconstructed from presentation-after attributes TEI apparatus handling: <lem> (main reading) preserved… See the full description on the dataset page: https://huggingface.co/datasets/sol-r/historica-corpus.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes39downloads
Dataset Card

Historica Corpus v2

A monolingual pretraining corpus of 314,438 passages (177M words, 1.1B characters) spanning 15 ancient and historical languages, from 500 BCE to 1750 CE.

What's New in v2

  • 2x more passages (314k vs 159k) due to varied-length chunking
  • SGML entity resolution: Middle English þ/ȝ/ð properly rendered (928k entities fixed)
  • PROIEL punctuation: reconstructed from presentation-after attributes
  • TEI apparatus handling: <lem> (main reading) preserved, <rdg> (variants) skipped, <supplied> kept
  • OCR artifact cleaning: line-break hyphenation, column numbers stripped
  • Varied chunk lengths: 11% short (100-200w), 24% medium (200-400w), 32% mid (400-700w), 23% long (700-1000w), 10% full (1000-1200w)
  • Metadata fixes: saga se → Swedish (not Sami), Coptic not hardcoded as Christian, OE genre not hardcoded as poetry

Languages

LanguageCodePassagesWordsSources
Latinlat196,632125MPL, CSEL, CAMENA, First1KGreek, Tesserae, Latin Library, CroALa, Corpus Iuris
Ancient Greekgrc42,32432MFirst1KGreek, PROIEL
Englisheng39,26115Menglish_trans, Corpus Iuris
Middle Englishenm27,8684MMichigan ME Corpus
Old Norsenon3,3160.6MSagaDB, CLTK, Heimskringla
Copticcop3,1050.4MCoptic Scriptorium
Germandeu456First1KGreek translations
Old Englishang373OE Sacred, OEDT
Norwegian (Bokmål)nob352SagaDB translations
Swedishswe219SagaDB translations
Frenchfra206SagaDB translations
Gothicgot134PROIEL (Wulfila Bible)
Old Church Slavonicchu85PROIEL (Codex Marianus)
Danishdan74SagaDB translations
Classical Armenianxcl33PROIEL

Sources

SourcePassagesDescription
CAMENA84,036Neo-Latin literature 1500-1750 (letters, history, poetry, encyclopedias)
Patrologia Latina60,403Church fathers (Latin, 4th-13th c.)
First1KGreek44,884Greek literature 700 BCE-900 CE
english_trans29,889English translations of classical texts (long-s OCR corrected)
Middle English27,868Michigan corpus (SGML entities resolved to Unicode)
Latin Library24,952Classical and medieval Latin
Tesserae10,896Classical Latin (intertextuality project)
CSEL10,556Church fathers (critical editions)
Corpus Iuris9,490Roman law (Latin + English)
SagaDB3,862Old Norse sagas + translations
Coptic Scriptorium3,102Coptic texts
Old Norse CLTK1,631Old Norse poetry + prose
PROIEL1,505Parallel treebank (with reconstructed punctuation)
CroALa919Croatian Latin
Others1,340OE Sacred, OEDT, Heimskringla

Schema

ColumnTypeDescription
sourcestringSource repository/collection
languagestringISO 639-3 code
authorstringAuthor (where known)
workstringWork title
genrestringGenre (poetry, history, law, etc.)
traditionstringTradition (christian, secular, norse_pagan)
urnstringCTS/URN identifier (where available)
idstringUnique passage identifier
textstringPassage text (cleaned, entity-resolved)
word_countintWord count
char_countintCharacter count

Extraction

Built with extract_corpus.py using a parser class architecture:

  • TEIParser — CSEL, PL, CAMENA, First1KGreek, CroALa, Coptic (with proper <lem>/<supplied> handling)
  • EnglishTransParser — TEI + long-s OCR correction
  • ProielParser — punctuation reconstructed from presentation-after
  • SagaParser — entity resolution, corrected language codes
  • MiddleEnglishParser — 203 SGML entities resolved to Unicode
  • PlaintextParser, TesseraeParser, OldEnglishParser