datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
between-the-line
暗流 Between the Line
暗流(Between the Line) 是一个中文攻击性 / 毒性短文本标注数据集,用于训练与评测「守望」微博理性发言主动干预插件的检测模型(望潮 TideWatcher)。名字取"字里行间"之意:攻击与恶意往往藏在字里行间。
1. 解决什么问题
中文互联网的攻击性发言检测缺少高质量的短句标注数据。公开的冒犯性语料覆盖有限,对缩写攻击(nmsl/cnm)、谐音、阴阳怪气、玩梗边界、冷暴力句式等"表达代差"覆盖不足。暗流以微博语境的短句为主,补充这些表达变体,为攻击性文本检测模型的训练与评测提供标注数据。
2. 主要内容
数据集共 2,414 条(train 1,936 + test 239 + val 239),本次发布 train(1,936 条)、test(239 条)与 val(239 条) 三个部分。
文件
条数
说明
train.csv
1,936
训练集(label 1: 976 / 0: 960,含 387… See the full description on the dataset page: https://huggingface.co/datasets/RainbowLIght/between-the-line.Analyst-Report-Biasweather-aus-rain-datasetirisRainPredictionVeriTrip
