CoolFace
Datasetpublic

willem640/MAAT_documentary_with_DDbDP_and_HGV_metadata

This is dataset containing the documentary papyri from MAAT, with metadata from the HGV and DDbDP. The scripts in preprocessing/maat were used to generate these files. The processing is not perfect, but this should have metadata for most papyri where it was available. The fields corpus_id, file_id, block_index, id, title, material, language, training_text and test_cases were derived from MAAT: Fitzgerald, W., & Barney, J. (2024). The Machine-Actionable Ancient Text (MAAT) Corpus (1.0.0-beta)… See the full description on the dataset page: https://huggingface.co/datasets/willem640/MAAT_documentary_with_DDbDP_and_HGV_metadata.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes17downloads
Dataset Card

dataset_info: features:

  • —name: corpus_id dtype: string
  • —name: file_id dtype: string
  • —name: block_index dtype: int64
  • —name: id dtype: string
  • —name: title dtype: string
  • —name: material dtype: string
  • —name: language dtype: string
  • —name: training_text dtype: string
  • —name: test_cases list:
  • —name: case_index dtype: int64
  • —name: id dtype: string
  • —name: test_case dtype: string
  • —name: hgv_id dtype: string
  • —name: tm_id dtype: string
  • —name: hgv_title dtype: string
  • —name: orig_place struct:
  • —name: PL dtype: string
  • —name: TM dtype: string
  • —name: text dtype: string
  • —name: text_classes list: string splits:
  • —name: train numbytes: 1141458209 numexamples: 56963 downloadsize: 143263318 datasetsize: 1141458209 configs:
  • —configname: default datafiles:
  • —split: train path: data/train-* license: cc-by-4.0 ---

This is dataset containing the documentary papyri from MAAT, with metadata from the HGV and DDbDP. The scripts in preprocessing/maat were used to generate these files. The processing is not perfect, but this should have metadata for most papyri where it was available.

The fields corpusid, fileid, blockindex, id, title, material, language, trainingtext and test_cases were derived from MAAT:

Fitzgerald, W., & Barney, J. (2024). The Machine-Actionable Ancient Text (MAAT) Corpus (1.0.0-beta) [Data set]. Machine Learning for Ancient Languages, ACL 2024 Workshop (ML3AL), Hybrid in Bangkok, Thailand and remote. WMU Herculaneum Project. https://doi.org/10.5281/zenodo.12553283. MAAT is licensed under the Creative Commons Attribution 4.0 License

The other fields are derived from the papyri.info project data. The tmid and hgvid fields come fro the Duke databank of Documentary Papyri and the hgvtitle, textclasses, daterange and origplace fields are derived from the Heidelberger Gesamtverzeichnis der griechischen Papyrusurkunden Ägyptens. Their data was made available under a Creative Commons Attribution 3.0 License, with copyright and attribution to the respective projects.

This data is available under the Creative Commons Attribution 4.0 License, with copyright and attribution to the original projects.