Weaxs/csc
Dataset for CSC 中文纠错数据集 Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. 共计 120w 条数据,以下是数据来源 数据集 语料 链接 SIGHAN+Wang271K 拼写纠错数据集 SIGHAN+Wang271K(27万条) https://huggingface.co/datasets/shibing624/CSC ECSpell 拼写纠错数据集 包含法律、医疗、金融等领域 https://github.com/Aopolin-Lv/ECSpell CGED 语法纠错数据集 仅包含了2016和2021年的数据集… See the full description on the dataset page: https://huggingface.co/datasets/Weaxs/csc.
7140
Dataset for CSC
中文纠错数据集
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
共计 120w 条数据,以下是数据来源
其余的数据集还可以看
- 中文文本纠错数据集汇总 (天池):https://tianchi.aliyun.com/dataset/138195
- NLPCC 2023中文语法纠错数据集:http://tcci.ccf.org.cn/conference/2023/taskdata.php
Languages
The data in CSC are in Chinese.
Dataset Structure
An example of "train" looks as follows:
{
"conversations": [
{"from":"human","value":"对这个句子纠错\n\n以后,我一直以来自学汉语了。"},
{"from":"gpt","value":"从此以后,我就一直自学汉语了。"}
]
}
Contributions
Weaxs 整理并上传
