CoolFace
Datasetpublic

UTSNLPGroup/PCR-ToxiCN

PCR-ToxiCN PCR-ToxiCN is a 500-example Chinese dataset for testing how well models spot offensive language hidden by phonetic cloaking (homophones and near-homophones). Field Type Notes text string Original Xiaohongshu comment offensive_label int 1 = offensive, 0 = non-offensive (250 / 250) strategy string HR, AR, NR, or MR Strategy What it is Example HR Hanzi replacement “沸物” → “废物” AR Alphabet / pinyin “SB” → “傻逼” NR Numerals as sounds “4”… See the full description on the dataset page: https://huggingface.co/datasets/UTSNLPGroup/PCR-ToxiCN.

sourceHugging Faceupdated 1y agoView on Hugging Face
6likes71downloads
10 commits on main
eb6839b1y ago

Update README.md

Hongbin37
bf0f7791y ago

Delete PCR-ToixCN.json

Hongbin37
0c8fd8d1y ago

Upload PCR-ToxiCN.json

Hongbin37
73367a31y ago

Update README.md

Hongbin37
6aa583f1y ago

Update README.md

Hongbin37
5afc1bc1y ago

Update README.md

Hongbin37
1664ba11y ago

Update README.md

Hongbin37
11bf1781y ago

Create README.md

Hongbin37
b52a0f81y ago

Upload PCR-ToixCN.json

Hongbin37
cd07a2c1y ago

initial commit

Hongbin37