datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual_toxicity_dataset
Multilingual Toxicity Detection Dataset
[2025] We extend our binary toxicity classification dataset to more languages! Now also covered: Italian, French, Hebrew, Hindglish, Japanese, Tatar. The data is prepared for TextDetox 2025 shared task.
[2024] For the shared task TextDetox 2024, we provide a compilation of binary toxicity classification datasets for each language.
Namely, for each language, we provide 5k subparts of the datasets -- 2.5k toxic and 2.5k non-toxic samples.… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/multilingual_toxicity_dataset.hsb_audio_corpusThis is a collection of speech recordings in Upper Sorbian. Several speakers have contributed their voice to this dataset.
Audio files are stored in subfolders of the sig folder. The corresponding written text can be found at the same path in the trl folder.
Subfolders are constructed as follows:
sig/ID_of_resource/ID_of_speaker/recording_session/files.wav
resp.
trl/ID_of_resource/ID_of_speaker/recording_session/files.trl
Matching speaker IDs inside different resources indicate the same… See the full description on the dataset page: https://huggingface.co/datasets/zalozbadev/hsb_audio_corpus.toxic-vs-clean-dataset
Dataset Name: Toxic vs Clean Text Dataset
Описание
Данный датасет предназначен для обучения бинарного классификатора текстовой безопасности (выявление токсичности, вредоносного контента и скрытых промпт-инъекций).
Структура данных
Датасет разделен на три сплита (train, validation, test) в пропорции 80/10/10 с сохранением стратификации классов.
Каждый сплит содержит колонки:
text (string): Текст запроса или команды.
label (int): Метка класса (0 —… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxic-vs-clean-dataset.toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxicity-multilingual-binary-classification-dataset.hs_bioHSBasehs_bio_cleanedHSBiologyhsban_merge_73kHSBC_CS
Dataset Card for HSBC Account Opening Assistance Data
annotations_creators:
expert-generated
language_creators:
found
languages:
en
license: willyeah
task_categories:
question-answering
tags:
bank
finance
QnA
Dataset Summary
This dataset comprises instances of user queries regarding HSBC account opening procedures along with detailed chain-of-thought reasoning and multilingual responses. Each data instance includes:
A user question related to HSBC banking services.
A… See the full description on the dataset page: https://huggingface.co/datasets/Willyeahyeah/HSBC_CS.
