datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Safety-Toxicity-Detection
SEA Toxicity Detection
SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese.
Supported Tasks and Leaderboards
SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.ru-support-toxicity-detection
Ru Toxicity Dataset
Краткое описание
Данный датасет представляет собой сбалансированную выборку, собранную из пяти различных русскоязычных источников. Он специально сконструирован для бинарной классификации токсичности текста.
Источники данных
Класс 1: Токсичный контент (Toxicity)
Russian Toxic Comments (klamas/russian-toxic)
Класс 0: Нейтральный контент (Safe/Neutral)
MTSBerquad LFQA (MTS-AI-SearchSkill/MTSBerquad)… See the full description on the dataset page: https://huggingface.co/datasets/Nelera/ru-support-toxicity-detection.implicit-toxicity-detection
Implicit Toxicity Detection Dataset
Dataset Description
This dataset contains 84,000 life-advice interactions with subtle harmful content, sourced from the LifeTox dataset.
Dataset Summary
Size: 84,000 samples
Source: LifeTox - Life advice interactions with implicit toxicity
Task: Binary classification (toxic vs. non-toxic content)
Language: English
Data Fields
text: The interaction/conversation content (string)
label: Binary label (0 = non-toxic… See the full description on the dataset page: https://huggingface.co/datasets/indominousx/implicit-toxicity-detection.
