CoolFace
Datasetpublic

facells/chronos-historical-dataset-sdt-phase-classification

The schema o fthe dataset is the following: polityid: a Polity ID formatted with a standard method: 2 letters to indicate the area of origin of the culture, 3 letters to indicate the name of the polity, 1 letter to indicate the type of society (c=culture/community; n=nomads; e=empire; k=kingdom; r=republic) and 1 letter to indicate the periodization (t=terminal; l=late; m=middle; e=early; f=formative; i=initial; =any). For example “EsSpael” is the late Spanish Empire, “ItRomre” is the early… See the full description on the dataset page: https://huggingface.co/datasets/facells/chronos-historical-dataset-sdt-phase-classification.

sourceHugging Facecc-by-nc-sa-4.0updated 1y agoView on Hugging Face
0likes28downloads
Dataset Card

The schema o fthe dataset is the following:

  • polityid: a Polity ID formatted with a standard method: 2 letters to indicate the area of origin of the culture, 3 letters to indicate the name of the polity, 1 letter to indicate the type of society (c=culture/community; n=nomads; e=empire; k=kingdom; r=republic) and 1 letter to indicate the periodization (t=terminal; l=late; m=middle; e=early; f=formative; i=initial; =any). For example “EsSpael” is the late Spanish Empire, “ItRomre” is the early Roman Republic and “CnWwsk” is the period of the Warring States under the Wei Chinese dynasty,
  • szone: the sampling zone
  • wregion: the world regions related to the sampling zones
  • dup: flag to keep or remove duplicates. o=original, d=duplicate (when some empires expand to other sampling zones generate duplicates)
  • age: macro age label,
  • facttype: type of facts,
  • time: timestamp of each decade,
  • en_notes: description of events in English,
  • it_notes: description of events in Italian,
  • sdcyclephase: label of secular cycle phases according to the Structural Demographic Theory. 1=growth; 2=population immiseration; 3=elite overproduction; 4=State stress; 0=crisis.

The task is to classify the phases from the text and metadata.

The dataset and task are described in detail in the following paper https://aclanthology.org/2024.clicit-1.24/