datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.csc-wireless-latency-synthetic-100k
CSC Wireless Latency Synthetic Dataset (100k)
This synthetic dataset provides 100,000 prompt-completion pairs designed for training and evaluating PHY/MAC cross-layer optimization models in hybrid Li-Fi/RF wireless networks.
Official Core Implementation & Runtime
To parse, simulate, or process this dataset according to the official protocol specifications, please utilize the official runtime library:
Core Protocol Library (npm):… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-wireless-latency-synthetic-100k.CSC-gpt4
Dataset Card for Chinese Spelling Correction(gpt4 fixed version)
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC-gpt4.CSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/henryy1990/CSC.csc_public_de3
csc_public_de3数据集
数据来源
1.由人民日报/学习强国/chinese-poetry等高质量数据人工生成;
2.来自人民日报高质量语料;
3.来自学习强国网站的高质量语料;
4.源数据为qwen生成的好词好句;
5.古诗词chinese-poetry; 文言文garychowcmu/daizhigev20;
数据简介
该数据主要为'的地得'纠错;
其中训练数据130753条, 验证数据5545条, 测试数据5545条;
句子平均长度为36, 最长句子长度为414, 最短为5, 95%的为89, 75%的为46, 60%的为34;
每个句子中字的平均错误数为2;
数据详情
################################################################################################################################
train.json
130753… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_public_de3.csc_dataset
