Weaxs/csc
Dataset for CSC 中文纠错数据集 Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. 共计 120w 条数据,以下是数据来源 数据集 语料 链接 SIGHAN+Wang271K 拼写纠错数据集 SIGHAN+Wang271K(27万条) https://huggingface.co/datasets/shibing624/CSC ECSpell 拼写纠错数据集 包含法律、医疗、金融等领域 https://github.com/Aopolin-Lv/ECSpell CGED 语法纠错数据集 仅包含了2016和2021年的数据集… See the full description on the dataset page: https://huggingface.co/datasets/Weaxs/csc.
update
upload backup
Delete grammar
Delete SIGHAN+Wang271K
Delete NLPCC
Delete ECSpell
Delete CGED
Delete NLG
Update .gitattributes
Update README.md
Update README.md
Upload 3 files
Rename validate.jsonl to validation.jsonl
Upload validate.jsonl
Update README.md
Update README.md
Update README.md
gpt merged dataset
del nlpcc18
upload dataset
Upload 3 files
Upload 4 files
initial commit
