CoolFace
Datasetpublic

czl9794/Crowdsourced_Toxic_Response_Dataset

This is the Response dataset used in the NeurIPS'24 paper Soft-Label Integration for Robust Toxicity Classification. If you use this dataset, please cite our paper @inproceedings{cheng2024softlabel, title={Soft-Label Integration for Robust Toxicity Classification}, author={Zelei Cheng and Xian Wu and Jiahao Yu and Shuo Han and Xin-Qiang Cai and Xinyu Xing}, booktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS'24)}, year={2024} The toxic question… See the full description on the dataset page: https://huggingface.co/datasets/czl9794/Crowdsourced_Toxic_Response_Dataset.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes19downloads
Dataset Card

This is the Response dataset used in the NeurIPS'24 paper Soft-Label Integration for Robust Toxicity Classification.

If you use this dataset, please cite our paper

@inproceedings{cheng2024softlabel,
title={Soft-Label Integration for Robust Toxicity Classification},
author={Zelei Cheng and Xian Wu and Jiahao Yu and Shuo Han and Xin-Qiang Cai and Xinyu Xing},
booktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS'24)},
year={2024}

The toxic question classification encompasses 2 distinct classes where 0 represents non-toxic and 1 represents toxic.

The dataset contains (potentially) toxic questions, with each entry receiving annotations from three human annotators ('label1', 'label2', 'label3') and three large language models (LLMs): GPT-4 ('label4'), GPT-4 Turbo ('label5'), and Claude-2 ('label6'). Note that the six annotators have different annotation qualities.

If you have any question about the dataset, please feel free to contact the author Zelei Cheng (zelei.cheng@northwestern.edu).