CoolFace
Datasetpublic

UTSNLPGroup/PCR-ToxiCN

PCR-ToxiCN PCR-ToxiCN is a 500-example Chinese dataset for testing how well models spot offensive language hidden by phonetic cloaking (homophones and near-homophones). Field Type Notes text string Original Xiaohongshu comment offensive_label int 1 = offensive, 0 = non-offensive (250 / 250) strategy string HR, AR, NR, or MR Strategy What it is Example HR Hanzi replacement “沸物” → “废物” AR Alphabet / pinyin “SB” → “傻逼” NR Numerals as sounds “4”… See the full description on the dataset page: https://huggingface.co/datasets/UTSNLPGroup/PCR-ToxiCN.

sourceHugging Faceupdated 1y agoView on Hugging Face
6likes71downloads

UTSNLPGroup/PCR-ToxiCN · main · files are served by the source, never re-hosted here