ukr-detect/ukr-toxicity-dataset-seminatural
Ukrainian Toxicity Dataset (Semi-natural) This is the first of its kind toxicity classification dataset for the Ukrainian language. The datasets was obtained semi-automatically by toxic keywords filtering. For manually collected datasets with crowdsourcing, please, check textdetox/multilingual_toxicity_dataset. Due to the subjective nature of toxicity, definitions of toxic language will vary. We include items that are commonly referred to as vulgar or profane language. (NLLB… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-toxicity-dataset-seminatural.
This repository belongs to ukr-detect on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
