mteb/InappropriatenessClassificationv2
InappropriatenessClassificationv2 An MTEB dataset Massive Text Embedding Benchmark Inappropriateness identification in the form of binary classification Task category t2t Domains Web, Social, Written Reference https://aclanthology.org/2021.bsnlp-1.4 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["InappropriatenessClassificationv2"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/InappropriatenessClassificationv2.
021
1---2annotations_creators:3- human-annotated4language:5- rus6license: cc-by-nc-sa-4.07multilinguality: monolingual8task_categories:9- text-classification10task_ids:11- sentiment-analysis12- sentiment-scoring13- sentiment-classification14- hate-speech-detection15dataset_info:16 features:17 - name: text18 dtype: string19 - name: label20 dtype: int6421 splits:22 - name: train23 num_bytes: 58124624 num_examples: 300025 - name: test26 num_bytes: 58138827 num_examples: 300028 download_size: 65076529 dataset_size: 116263430configs:31- config_name: default32 data_files:33 - split: train34 path: data/train-*35 - split: test36 path: data/test-*37tags:38- mteb39- text40---41<!-- adapted from https://github.com/huggingface/huggingface_hub/blob/v0.30.2/src/huggingface_hub/templates/datasetcard_template.md -->42 43<div align="center" style="padding: 40px 20px; background-color: white; border-radius: 12px; box-shadow: 0 2px 10px rgba(0, 0, 0, 0.05); max-width: 600px; margin: 0 auto;">44 <h1 style="font-size: 3.5rem; color: #1a1a1a; margin: 0 0 20px 0; letter-spacing: 2px; font-weight: 700;">InappropriatenessClassificationv2</h1>45 <div style="font-size: 1.5rem; color: #4a4a4a; margin-bottom: 5px; font-weight: 300;">An <a href="https://github.com/embeddings-benchmark/mteb" style="color: #2c5282; font-weight: 600; text-decoration: none;" onmouseover="this.style.textDecoration='underline'" onmouseout="this.style.textDecoration='none'">MTEB</a> dataset</div>46 <div style="font-size: 0.9rem; color: #2c5282; margin-top: 10px;">Massive Text Embedding Benchmark</div>47</div>48 49Inappropriateness identification in the form of binary classification50 51| | |52|---------------|---------------------------------------------|53| Task category | t2t |54| Domains | Web, Social, Written |55| Reference | https://aclanthology.org/2021.bsnlp-1.4 |56 57 58## How to evaluate on this task59 60You can evaluate an embedding model on this dataset using the following code:61 62```python63import mteb64 65task = mteb.get_tasks(["InappropriatenessClassificationv2"])66evaluator = mteb.MTEB(task)67 68model = mteb.get_model(YOUR_MODEL)69evaluator.run(model)70```71 72<!-- Datasets want link to arxiv in readme to autolink dataset with paper -->73To learn more about how to run models on `mteb` task check out the [GitHub repitory](https://github.com/embeddings-benchmark/mteb). 74 75## Citation76 77If you use this dataset, please cite the dataset as well as [mteb](https://github.com/embeddings-benchmark/mteb), as this dataset likely includes additional processing as a part of the [MMTEB Contribution](https://github.com/embeddings-benchmark/mteb/tree/main/docs/mmteb).78 79```bibtex80 81@inproceedings{babakov-etal-2021-detecting,82 abstract = {Not all topics are equally {``}flammable{''} in terms of toxicity: a calm discussion of turtles or fishing less often fuels inappropriate toxic dialogues than a discussion of politics or sexual minorities. We define a set of sensitive topics that can yield inappropriate and toxic messages and describe the methodology of collecting and labelling a dataset for appropriateness. While toxicity in user-generated data is well-studied, we aim at defining a more fine-grained notion of inappropriateness. The core of inappropriateness is that it can harm the reputation of a speaker. This is different from toxicity in two respects: (i) inappropriateness is topic-related, and (ii) inappropriate message is not toxic but still unacceptable. We collect and release two datasets for Russian: a topic-labelled dataset and an appropriateness-labelled dataset. We also release pre-trained classification models trained on this data.},83 address = {Kiyv, Ukraine},84 author = {Babakov, Nikolay and85Logacheva, Varvara and86Kozlova, Olga and87Semenov, Nikita and88Panchenko, Alexander},89 booktitle = {Proceedings of the 8th Workshop on Balto-Slavic Natural Language Processing},90 editor = {Babych, Bogdan and91Kanishcheva, Olga and92Nakov, Preslav and93Piskorski, Jakub and94Pivovarova, Lidia and95Starko, Vasyl and96Steinberger, Josef and97Yangarber, Roman and98Marci{\'n}czuk, Micha{\l} and99Pollak, Senja and100P{\v{r}}ib{\'a}{\v{n}}, Pavel and101Robnik-{\v{S}}ikonja, Marko},102 month = apr,103 pages = {26--36},104 publisher = {Association for Computational Linguistics},105 title = {Detecting Inappropriate Messages on Sensitive Topics that Could Harm a Company{'}s Reputation},106 url = {https://aclanthology.org/2021.bsnlp-1.4},107 year = {2021},108}109 110 111@article{enevoldsen2025mmtebmassivemultilingualtext,112 title={MMTEB: Massive Multilingual Text Embedding Benchmark},113 author={Kenneth Enevoldsen and Isaac Chung and Imene Kerboua and Márton Kardos and Ashwin Mathur and David Stap and Jay Gala and Wissam Siblini and Dominik Krzemiński and Genta Indra Winata and Saba Sturua and Saiteja Utpala and Mathieu Ciancone and Marion Schaeffer and Gabriel Sequeira and Diganta Misra and Shreeya Dhakal and Jonathan Rystrøm and Roman Solomatin and Ömer Çağatan and Akash Kundu and Martin Bernstorff and Shitao Xiao and Akshita Sukhlecha and Bhavish Pahwa and Rafał Poświata and Kranthi Kiran GV and Shawon Ashraf and Daniel Auras and Björn Plüster and Jan Philipp Harries and Loïc Magne and Isabelle Mohr and Mariya Hendriksen and Dawei Zhu and Hippolyte Gisserot-Boukhlef and Tom Aarsen and Jan Kostkan and Konrad Wojtasik and Taemin Lee and Marek Šuppa and Crystina Zhang and Roberta Rocca and Mohammed Hamdy and Andrianos Michail and John Yang and Manuel Faysse and Aleksei Vatolin and Nandan Thakur and Manan Dey and Dipam Vasani and Pranjal Chitale and Simone Tedeschi and Nguyen Tai and Artem Snegirev and Michael Günther and Mengzhou Xia and Weijia Shi and Xing Han Lù and Jordan Clive and Gayatri Krishnakumar and Anna Maksimova and Silvan Wehrli and Maria Tikhonova and Henil Panchal and Aleksandr Abramov and Malte Ostendorff and Zheng Liu and Simon Clematide and Lester James Miranda and Alena Fenogenova and Guangyu Song and Ruqiya Bin Safi and Wen-Ding Li and Alessia Borghini and Federico Cassano and Hongjin Su and Jimmy Lin and Howard Yen and Lasse Hansen and Sara Hooker and Chenghao Xiao and Vaibhav Adlakha and Orion Weller and Siva Reddy and Niklas Muennighoff},114 publisher = {arXiv},115 journal={arXiv preprint arXiv:2502.13595},116 year={2025},117 url={https://arxiv.org/abs/2502.13595},118 doi = {10.48550/arXiv.2502.13595},119}120 121@article{muennighoff2022mteb,122 author = {Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo{\"\i}c and Reimers, Nils},123 title = {MTEB: Massive Text Embedding Benchmark},124 publisher = {arXiv},125 journal={arXiv preprint arXiv:2210.07316},126 year = {2022}127 url = {https://arxiv.org/abs/2210.07316},128 doi = {10.48550/ARXIV.2210.07316},129}130```131 132# Dataset Statistics133<details>134 <summary> Dataset Statistics</summary>135 136The following code contains the descriptive statistics from the task. These can also be obtained using:137 138```python139import mteb140 141task = mteb.get_task("InappropriatenessClassificationv2")142 143desc_stats = task.metadata.descriptive_stats144```145 146```json147{148 "test": {149 "num_samples": 3000,150 "number_of_characters": 304259,151 "number_texts_intersect_with_train": 0,152 "min_text_length": 15,153 "average_text_length": 101.41966666666667,154 "max_text_length": 2159,155 "unique_texts": 3000,156 "min_labels_per_text": 1,157 "average_label_per_text": 1.0,158 "max_labels_per_text": 1,159 "unique_labels": 2,160 "labels": {161 "0": {162 "count": 2110163 },164 "1": {165 "count": 890166 }167 }168 },169 "train": {170 "num_samples": 3000,171 "number_of_characters": 304641,172 "number_texts_intersect_with_train": null,173 "min_text_length": 19,174 "average_text_length": 101.547,175 "max_text_length": 1802,176 "unique_texts": 3000,177 "min_labels_per_text": 1,178 "average_label_per_text": 1.0,179 "max_labels_per_text": 1,180 "unique_labels": 2,181 "labels": {182 "0": {183 "count": 2126184 },185 "1": {186 "count": 874187 }188 }189 }190}191```192 193</details>194 195---196*This dataset card was automatically generated using [MTEB](https://github.com/embeddings-benchmark/mteb)*