CoolFace
Datasetpublic

AncientLanguages/Latin-CC-170M

Latin-CC-170M Reupload of the Corpus Corporum as Parquet, originally from Kaggle, and reuploaded as CSV by Fece228/latin-literature-dataset-170M. License These works are public domain. Original README This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes: Classical Latin: works of Caesar, Cicero and many more Medieval… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.

sourceHugging Faceupdated 9mo agoView on Hugging Face
1likes56downloads
Dataset Card

Latin-CC-170M

Reupload of the Corpus Corporum as Parquet, originally from Kaggle, and reuploaded as CSV by Fece228/latin-literature-dataset-170M.

License

These works are public domain.


Original README

This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes:

  • Classical Latin: works of Caesar, Cicero and many more
  • Medieval Latin: a substantial amount of religious texts by Thomas Aquinas, Bonaventura and others
  • Phliosophical works written by Descartes, Spinosa, …
  • Regional Latin literature of Croatian, German, Italian authors