datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
people_relation_classification本数据集用于人物关系分类,一共14种关系类型:不确定, 夫妻, 父母, 兄弟姐妹, 上下级, 师生, 好友, 同学, 合作, 同一个人, 情侣, 祖孙, 同门, 亲戚。
本数据集共3881条,其中训练集3105条,测试集776条,参看train.csv和test.csv。
数据集的人物关系分布如下:
关于使用R-BERT模型训练该数据集,可参考文章:NLP(四十二)人物关系分类的再次尝试.
Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset.
https://arxiv.org/abs/2311.06838
Cite:
@article{gan2023giellm,
title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2311.06838},
year={2023}
}
The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html
