BAAI/CCI3-HQ-Annotation-Benchmark
CCI3-HQ-Annotation-Benchmark These 14k samples were randomly extracted from a large corpus of Chinese texts, containing both the original text and corresponding labels. They can be used to evaluate the quality of Chinese corpora. Citation If you use this benchmark or the CCI3-HQ dataset, please cite: @misc{wang2024cci30hqlargescalechinesedataset, title={CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CCI3-HQ-Annotation-Benchmark.
CCI3-HQ-Annotation-Benchmark
These 14k samples were randomly extracted from a large corpus of Chinese texts, containing both the original text and corresponding labels. They can be used to evaluate the quality of Chinese corpora.
Citation
If you use this benchmark or the CCI3-HQ dataset, please cite:
@misc{wang2024cci30hqlargescalechinesedataset,
title={CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models},
author={Liangdong Wang and Bo-Wen Zhang and Chengwei Wu and Hanyu Zhao and Xiaofeng Shi and Shuhao Gu and Jijie Li and Quanyue Ma and TengFei Pan and Guang Liu},
year={2024},
eprint={2410.18505},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.18505}
}