datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual_toxicity_dataset
Multilingual Toxicity Detection Dataset
[2025] We extend our binary toxicity classification dataset to more languages! Now also covered: Italian, French, Hebrew, Hindglish, Japanese, Tatar. The data is prepared for TextDetox 2025 shared task.
[2024] For the shared task TextDetox 2024, we provide a compilation of binary toxicity classification datasets for each language.
Namely, for each language, we provide 5k subparts of the datasets -- 2.5k toxic and 2.5k non-toxic samples.… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/multilingual_toxicity_dataset.AI-vs-Real
🖼️ AI-vs-Real Dataset
A balanced dataset for AI-generated vs Real image classification.This dataset is designed to help researchers, developers, and practitioners build and evaluate models that can distinguish between synthetic (AI-generated) and authentic (human-captured) images.
📊 Dataset Overview
Classes:
0 → AI-generated images
1 → Real (human-captured) images
Balance:The dataset is properly balanced across both classes.This ensures that… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/AI-vs-Real.toxic-vs-clean-dataset
Dataset Name: Toxic vs Clean Text Dataset
Описание
Данный датасет предназначен для обучения бинарного классификатора текстовой безопасности (выявление токсичности, вредоносного контента и скрытых промпт-инъекций).
Структура данных
Датасет разделен на три сплита (train, validation, test) в пропорции 80/10/10 с сохранением стратификации классов.
Каждый сплит содержит колонки:
text (string): Текст запроса или команды.
label (int): Метка класса (0 —… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxic-vs-clean-dataset.toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxicity-multilingual-binary-classification-dataset.hs_biohs_bio_cleaned
