willem640/MAAT_documentary_with_DDbDP_and_HGV_metadata
This is dataset containing the documentary papyri from MAAT, with metadata from the HGV and DDbDP. The scripts in preprocessing/maat were used to generate these files. The processing is not perfect, but this should have metadata for most papyri where it was available. The fields corpus_id, file_id, block_index, id, title, material, language, training_text and test_cases were derived from MAAT: Fitzgerald, W., & Barney, J. (2024). The Machine-Actionable Ancient Text (MAAT) Corpus (1.0.0-beta)… See the full description on the dataset page: https://huggingface.co/datasets/willem640/MAAT_documentary_with_DDbDP_and_HGV_metadata.
dataset_info: features:
- name: corpus_id dtype: string
- name: file_id dtype: string
- name: block_index dtype: int64
- name: id dtype: string
- name: title dtype: string
- name: material dtype: string
- name: language dtype: string
- name: training_text dtype: string
- name: test_cases list:
- name: case_index dtype: int64
- name: id dtype: string
- name: test_case dtype: string
- name: hgv_id dtype: string
- name: tm_id dtype: string
- name: hgv_title dtype: string
- name: orig_place struct:
- name: PL dtype: string
- name: TM dtype: string
- name: text dtype: string
- name: text_classes list: string splits:
- name: train numbytes: 1141458209 numexamples: 56963 downloadsize: 143263318 datasetsize: 1141458209 configs:
- configname: default datafiles:
- split: train path: data/train-* license: cc-by-4.0 ---
This is dataset containing the documentary papyri from MAAT, with metadata from the HGV and DDbDP. The scripts in preprocessing/maat were used to generate these files. The processing is not perfect, but this should have metadata for most papyri where it was available.
The fields corpusid, fileid, blockindex, id, title, material, language, trainingtext and test_cases were derived from MAAT:
Fitzgerald, W., & Barney, J. (2024). The Machine-Actionable Ancient Text (MAAT) Corpus (1.0.0-beta) [Data set]. Machine Learning for Ancient Languages, ACL 2024 Workshop (ML3AL), Hybrid in Bangkok, Thailand and remote. WMU Herculaneum Project. https://doi.org/10.5281/zenodo.12553283. MAAT is licensed under the Creative Commons Attribution 4.0 License
The other fields are derived from the papyri.info project data. The tmid and hgvid fields come fro the Duke databank of Documentary Papyri and the hgvtitle, textclasses, daterange and origplace fields are derived from the Heidelberger Gesamtverzeichnis der griechischen Papyrusurkunden Ägyptens. Their data was made available under a Creative Commons Attribution 3.0 License, with copyright and attribution to the respective projects.
This data is available under the Creative Commons Attribution 4.0 License, with copyright and attribution to the original projects.
