alkaid0618/ChineseWebText
ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText. ChineseWebText Dataset Overview We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is… See the full description on the dataset page: https://huggingface.co/datasets/alkaid0618/ChineseWebText.
0967
No commit history came back for main. The revision may not exist, or the source declined the request.
