CoolFace
Datasetpublic

facells/histohate-cultural-analytics-corpus

This is a synthetic corpus. We asked to Gemini 2.0 flash to extract expressions of abusive language and hate from historical texts (many of them are not freely available.) title: the title of the text (if any). Most english texts are anonymized, but all Italian titles are readable. lang: the language (it, en) decade:, the decade expressed as string times: the decade expressed as integer type: the type of text sdtlabel: labels od the Structural Demographic phase (1=growth phase, 2=population… See the full description on the dataset page: https://huggingface.co/datasets/facells/histohate-cultural-analytics-corpus.

sourceHugging Facecc-by-nc-sa-4.0updated 1y agoView on Hugging Face
0likes25downloads
Dataset Card

This is a synthetic corpus. We asked to Gemini 2.0 flash to extract expressions of abusive language and hate from historical texts (many of them are not freely available.)

title: the title of the text (if any). Most english texts are anonymized, but all Italian titles are readable.

lang: the language (it, en)

decade:, the decade expressed as string

times: the decade expressed as integer

type: the type of text

sdtlabel: labels od the Structural Demographic phase (1=growth phase, 2=population immiseration phase, 3=elite overproduction phase, 4=State stress phase, 5=crisis phase)

g2f score: abusive language score given by Gemini 2.0 flash, expressed as number

c35s score: abusive language score given by Claude 3.5 sonnet, expressed as number

g2f analysis: abusive text extracted by Gemini 2.0 flash and analysis.

The corpus has been used for mapping the abusive language scores to the SDT labels, finding that there is less abusive language during phases 2 of the secular cycles.