datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Safety-Toxicity-Detection
SEA Toxicity Detection
SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese.
Supported Tasks and Leaderboards
SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.ru-support-toxicity-detection
Ru Toxicity Dataset
Краткое описание
Данный датасет представляет собой сбалансированную выборку, собранную из пяти различных русскоязычных источников. Он специально сконструирован для бинарной классификации токсичности текста.
Источники данных
Класс 1: Токсичный контент (Toxicity)
Russian Toxic Comments (klamas/russian-toxic)
Класс 0: Нейтральный контент (Safe/Neutral)
MTSBerquad LFQA (MTS-AI-SearchSkill/MTSBerquad)… See the full description on the dataset page: https://huggingface.co/datasets/Nelera/ru-support-toxicity-detection.pharma-toxicity-decoherence-precursor-detection-v0.1What this dataset tests
Whether a system can detect early systemic decoherence that precedes overt toxicity.
It is not organ damage detection.
It is coupling loss detection.
Required outputs
decoherence_onset_time
decoherence_signature_set
affected_coupling_edges
precursor_pattern_label
time_to_overt_toxicity_estimate
confidence_score
Use case
Early safety screening.
Flag candidates that destabilize cross-tissue coherence
before late-stage attrition.
implicit-toxicity-detection
Implicit Toxicity Detection Dataset
Dataset Description
This dataset contains 84,000 life-advice interactions with subtle harmful content, sourced from the LifeTox dataset.
Dataset Summary
Size: 84,000 samples
Source: LifeTox - Life advice interactions with implicit toxicity
Task: Binary classification (toxic vs. non-toxic content)
Language: English
Data Fields
text: The interaction/conversation content (string)
label: Binary label (0 = non-toxic… See the full description on the dataset page: https://huggingface.co/datasets/indominousx/implicit-toxicity-detection.
