CoolFace
Datasetpublic

Toygar/turkish-offensive-language-detection

Dataset Summary This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.

sourceHugging Facecc-by-2.0updated 3y agoView on Hugging Face
20likes173downloads
9 commits on main
5aff92f3y ago

Update README.md

Toygar
b41782e4y ago

citation is added

Toygar
456e0e14y ago

Fix task_ids (#1)

Toygar, albertvillanova
990c7a54y ago

readme v0.2

Toygar
41707c84y ago

Add validation file

Toygar
4b72e8e4y ago

Add test file

Toygar
32e03514y ago

Add train file

Toygar
1d9303c4y ago

readme v0.1

Toygar
64798884y ago

initial commit

Toygar