CoolFace
Datasetpublic

yoonholee/homer-parallel-translations

Homer Parallel Translations Book-aligned parallel translations of Homer's Iliad and Odyssey by six and four English translators respectively, spanning four centuries of translation practice (1611-1900). All public domain in the US. Intended as a reference set for studying translator taste: each translator made distinct stylistic and interpretive choices on the same source text, and the dataset lets you compare them passage-by-passage. Contents 240 records = 24… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/homer-parallel-translations.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
0likes24downloads
Dataset Card

Homer Parallel Translations

Book-aligned parallel translations of Homer's Iliad and Odyssey by six and four English translators respectively, spanning four centuries of translation practice (1611-1900). All public domain in the US. Intended as a reference set for studying translator taste: each translator made distinct stylistic and interpretive choices on the same source text, and the dataset lets you compare them passage-by-passage.

Contents

240 records = 24 books × (6 Iliad + 4 Odyssey translators).

Iliad (144 records)

TranslatorYearStyle
George Chapman1611fourteeners (not included -- different book-marker convention)
Alexander Pope1720heroic couplets
William Cowper1791blank verse
Theodore Alois Buckley1851prose
Edward Smith-Stanley (Earl of Derby)1864blank verse
Lang, Leaf & Myers1883prose (literary archaism)
Samuel Butler1898prose (plain)

Odyssey (96 records)

TranslatorYearStyle
Alexander Pope1725heroic couplets
William Cowper1791blank verse
Butcher & Lang1879prose
Samuel Butler1900prose

Schema

ColumnTypeDescription
workstring"Iliad" or "Odyssey"
authorstring"Homer"
book_numberintBook number, 1-24
translatorstringTranslator name
yearintApproximate publication year
stylestringVerse / prose style descriptor
gutenberg_idintProject Gutenberg book ID
textstringFull text of the book in this translation
char_countintCharacter count
word_countintWord count

Why this exists

If you want to train or evaluate a model's taste for translation as interpretation, you need the same source text rendered by multiple expert hands. Modern translations of classical texts are almost all in copyright. Homer has the advantage that every major English translation from 1611 to 1900 is public domain, giving you a dense spectrum of translation philosophies on identical source material -- from Chapman's Elizabethan fourteeners through Pope's Augustan couplets to Butler's Victorian plain prose.

Use cases:

  • Style classification: given a passage, predict the translator.
  • Pairwise taste: "which rendering of this passage is more faithful? More beautiful? More readable today?"
  • Evaluation rubric development: what makes a great translation? Build eval components from dimensions where translators visibly disagree.
  • Few-shot prompting: given a new passage of Homer, generate a translation in the style of Pope, Butler, etc.

Source books

IDTitleTranslator
6130The IliadPope
16452The Iliad of HomerCowper
6150The IliadDerby
22382The IliadBuckley
3059The IliadLang, Leaf & Myers
2199The IliadButler
3160The OdysseyPope
24269The Odyssey of HomerCowper
1728The Odyssey of HomerButcher & Lang
1727The OdysseyButler

Parsing notes

Each translation was split on "BOOK I / II / ..." markers (or "BOOK THE FIRST / SECOND / ..." for Buckley). The first set of marker hits in each file corresponds to the table of contents, so we kept the last occurrence of each book number -- the one inside the body. Argument summaries, illustration tags, and footnote markers were stripped where regex-detectable; some editorial cruft remains. Chapman's 1611 Iliad uses a different book-marker convention and is not yet parsed into this release.

Book-level alignment means you can compare e.g. Butler's Book IX to Pope's Book IX, but passage-level alignment (individual lines or verses) would require a separate alignment pass -- the translations do not preserve line numbering and their structural choices differ.

License

CC0 1.0 Universal -- all source texts are public domain. Parsing and curation are also released under CC0.

Citation

Source texts courtesy of Project Gutenberg. Curation by Yoonho Lee.