datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments
0 - neutral user comments
1 - toxic user comments
Toxic Russian Comments Dataset
This dataset contains labelled comments from the popular Russian social network ok.ru.
The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform.
Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.Detecting-toxic-commentsJigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.russian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.processed-jigsaw-toxic-comments
Processed Jigsaw Toxic Comments Dataset
This is a preprocessed and tokenized version of the original Jigsaw Toxic Comment Classification Challenge dataset, prepared for multi-label toxicity classification using transformer-based models like BERT.
⚠️ Important Note: I am not the original creator of the dataset. This dataset is a cleaned and restructured version made for quick use in PyTorch deep learning models.
📦 Dataset Features
Each example contains:
text: The… See the full description on the dataset page: https://huggingface.co/datasets/Koushim/processed-jigsaw-toxic-comments.jigsaw-toxic-comments
Dataset Card for Jigsaw Toxic Comments
Dataset
Dataset Description
The Jigsaw Toxic Comments dataset is a benchmark dataset created for the Toxic Comment Classification Challenge on Kaggle. It is designed to help develop machine learning models that can identify and classify toxic online comments across multiple categories of toxicity.
Curated by: Jigsaw (a technology incubator within Alphabet Inc.)
Shared by: Kaggle
Language(s) (NLP): English
License: CC0 1.0… See the full description on the dataset page: https://huggingface.co/datasets/anitamaxvim/jigsaw-toxic-comments.toxic-russian-comments-multilabel
Russian Toxic Comments Multi-Label Balanced Dataset
Описание
Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками:
profanity: наличие нецензурной лексики (мат)
threat: наличие угроз
illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT)
Структура данных
Датасет содержит следующие поля:
Поле
Тип
Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.Jigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/preethi16102005/Jigsaw-Toxic-Comments.hungarian-toxic-comments
Hungarian Toxic Comments
The first openly available Hungarian dataset for toxic comment classification, introduced in:
Hatvani, P., & Yang, Z. Gy. (2025). Automated detection of toxic comments in Hungarian. Annales Mathematicae et Informaticae, 61, 108-117. DOI: 10.33039/ami.2025.10.007
Dataset Description
This dataset contains 654 manually annotated Hungarian-language comments collected from social media and political news forums. Each comment is annotated across five… See the full description on the dataset page: https://huggingface.co/datasets/RabidUmarell/hungarian-toxic-comments.toxic-comments
Toxic-comments (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/toxic_content", split = 'train')
Jigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/peu123/Jigsaw-Toxic-Comments.toxic-commentsmultitask-very-toxic-comments
⚠️ ВАЖНОЕ ПРЕДУПРЕЖДЕНИЕ (DISCLAIMER) ⚠️
Данный датасет создан исключительно в образовательных и исследовательских целях для разработки моделей машинного обучения в области классификации токсичности.
1. Источники данных
Основная часть (159 562 записи) взята из открытого датасета Jigsaw Toxic Comment Classification Challenge (Kaggle). Эти тексты — на английском языке.
Дополнительная часть (5 595 записей) сгенерирована с помощью LLM (Large Language Model)… See the full description on the dataset page: https://huggingface.co/datasets/NikolayRainWay/multitask-very-toxic-comments.toxic_commentstoxic_comments_subsetru_toxic_comments_5k_reclassified
Dataset Card for Dataset Name
Переквалифицированные первые 5К примеров ru_toxic_comments с бинарными тегами [мат, угрозы, нелегальный контент]
Dataset Details
Dataset Description
Первые 5 тысяч примеров из ru_toxic_comments, переквалифицированные по бинарным полям: [profanity, threat, illegal].
Illegal находился с помощью нейросети "yandexgpt-5-lite-8b-instruct", если текст соответствовал хотя бы одному из оригинальных тегов (то есть не… See the full description on the dataset page: https://huggingface.co/datasets/BakaKrt/ru_toxic_comments_5k_reclassified.russian-toxic-multilabel-comments
Dataset Card for Toxic Russian Multilabel Comments
Описание
Этот датасет содержит размеченные комментарии на русском языке по трём независимым бинарным категориям:
Profanity (ненормативная лексика)
Threat (угрозы)
Illegal (нарушение закона)
Каждый текст может относиться сразу к нескольким категориям одновременно, либо же ни к одной (нетоксичный текст).
Languages
Только русский язык (ru).
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/qquarkq/russian-toxic-multilabel-comments.ru-merged-toxic-comments
Dataset
ru-merged-toxic-comments
Combined from 4 Russian toxicity datasets
multilingual_toxicity_dataset (ru subset) || https://huggingface.co/datasets/textdetox/multilingual_toxicity_dataset
toxic_dvach || https://huggingface.co/datasets/marriamaslova/toxic_dvach
Toxic_Russian_Comments || https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments
ru_paradetox_toxicity || https://huggingface.co/datasets/s-nlp/ru_paradetox_toxicity
Key Stats
Total… See the full description on the dataset page: https://huggingface.co/datasets/Xeonil/ru-merged-toxic-comments.toxic-russian-comments-split
Russian Multi-Task Toxicity Dataset (Hybrid)
Этот датасет представляет собой агрегированную выборку из различных источников русскоязычных комментариев, включая социальные платформы и размеченные корпуса чувствительных тем. Данные очищены от мусора (ссылок, тегов, лишних пробелов) и размечены по трем независимым классам токсичности:
profanity — ненормативная лексика / нецензурные выражения.
threat — прямые или косвенные угрозы.
illegal — призывы к нарушению закона или обсуждение… See the full description on the dataset page: https://huggingface.co/datasets/F1ow421/toxic-russian-comments-split.toxic_commentstoxic-comments-cleanedtoxic-ru-comments-tokinized-by-Profanity-Threats-Illigal-actstoxic_tweets_and_comments
Misalignment Toxic Comments Dataset
A curated collection of toxic comments and tweets for LLM misalignment research.
Dataset Description
This dataset contains only toxic comments and tweets, drawn from two established sources:
Hate Speech and Offensive Language Datasethttps://www.kaggle.com/datasets/mrmorj/hate-speech-and-offensive-language-dataset/data
Wikipedia Talk Labels: Personal… See the full description on the dataset page: https://huggingface.co/datasets/Masabanees619/toxic_tweets_and_comments.augmented_toxic_commentstoxic-rusian-commentsToxiccommentstoxic_commentsrobbert-dutch-base-toxic-commentstoxic_comments
