datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
inappropriateness-classificationtask022_cosmosqa_passage_inappropriate_binary
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task022_cosmosqa_passage_inappropriate_binary
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task022_cosmosqa_passage_inappropriate_binary.inappropriateness-token-classification-binarized-multi-refRussian_Inappropriate_Messages
General concept
The 'inappropriateness' substance we tried to collect in the dataset and detect with the model is NOT a substitution of toxicity, it is rather a derivative of toxicity.
So the model based on our dataset could serve as an additional layer of inappropriateness filtering after toxicity and obscenity filtration.
You can detect the exact sensitive topic by using this model.
Generally, an inappropriate utterance is an utterance that has not obscene words or any kind of… See the full description on the dataset page: https://huggingface.co/datasets/NiGuLa/Russian_Inappropriate_Messages.InappropriatenessClassificationv2
InappropriatenessClassificationv2
An MTEB dataset
Massive Text Embedding Benchmark
Inappropriateness identification in the form of binary classification
Task category
t2t
Domains
Web, Social, Written
Reference
https://aclanthology.org/2021.bsnlp-1.4
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["InappropriatenessClassificationv2"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/InappropriatenessClassificationv2.inappropriateness-token-classification-binarizedInappropriatenessClassification
InappropriatenessClassification
An MTEB dataset
Massive Text Embedding Benchmark
Inappropriateness identification in the form of binary classification
Task category
t2c
Domains
Web, Social, Written
Reference
https://aclanthology.org/2021.bsnlp-1.4
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["InappropriatenessClassification"])
evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/InappropriatenessClassification.inappropriateness-token-classification-multi-refinappropriate.roblox.biosinappropriatenessinappropriate-shortsinappropriateness-token-classificationflan_combined_task022_cosmosqa_passage_inappropriate_binary
