CoolFace
Datasetpublic

Weaxs/csc

Dataset for CSC 中文纠错数据集 Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. 共计 120w 条数据,以下是数据来源 数据集 语料 链接 SIGHAN+Wang271K 拼写纠错数据集 SIGHAN+Wang271K(27万条) https://huggingface.co/datasets/shibing624/CSC ECSpell 拼写纠错数据集 包含法律、医疗、金融等领域 https://github.com/Aopolin-Lv/ECSpell CGED 语法纠错数据集 仅包含了2016和2021年的数据集… See the full description on the dataset page: https://huggingface.co/datasets/Weaxs/csc.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
7likes140downloads
Dataset Card

Dataset for CSC

中文纠错数据集

Dataset Description

Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.

共计 120w 条数据,以下是数据来源

数据集语料链接
SIGHAN+Wang271K 拼写纠错数据集SIGHAN+Wang271K(27万条)https://huggingface.co/datasets/shibing624/CSC
ECSpell 拼写纠错数据集包含法律、医疗、金融等领域https://github.com/Aopolin-Lv/ECSpell
CGED 语法纠错数据集仅包含了2016和2021年的数据集https://github.com/wdimmy/Automatic-Corpus-Generation?spm=a2c22.12282016.0.0.5f3e7398w7SL4P
NLPCC 纠错数据集包含语法纠错和拼写纠错https://github.com/Arvid-pku/NLPCC2023SharedTask8 <br/>http://tcci.ccf.org.cn/conference/2023/dldoc/nacgec_training.zip<br/>http://tcci.ccf.org.cn/conference/2018/dldoc/trainingdata02.tar.gz
pycorrector 语法纠错集中文语法纠错数据集https://github.com/shibing624/pycorrector/tree/llm/examples/data/grammar

其余的数据集还可以看

  • 中文文本纠错数据集汇总 (天池):https://tianchi.aliyun.com/dataset/138195
  • NLPCC 2023中文语法纠错数据集:http://tcci.ccf.org.cn/conference/2023/taskdata.php

Languages

The data in CSC are in Chinese.

Dataset Structure

An example of "train" looks as follows:

json
{
  "conversations": [
    {"from":"human","value":"对这个句子纠错\n\n以后,我一直以来自学汉语了。"},
    {"from":"gpt","value":"从此以后,我就一直自学汉语了。"}
  ]
}

Contributions

Weaxs 整理并上传