CoolFace
Datasetpublic

ukr-detect/ukr-toxicity-dataset-seminatural

Ukrainian Toxicity Dataset (Semi-natural) This is the first of its kind toxicity classification dataset for the Ukrainian language. The datasets was obtained semi-automatically by toxic keywords filtering. For manually collected datasets with crowdsourcing, please, check textdetox/multilingual_toxicity_dataset. Due to the subjective nature of toxicity, definitions of toxic language will vary. We include items that are commonly referred to as vulgar or profane language. (NLLB… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-toxicity-dataset-seminatural.

sourceHugging Faceopenrail++updated 2y agoView on Hugging Face
9likes95downloads

ukr-detect/ukr-toxicity-dataset-seminatural · main · files are served by the source, never re-hosted here