CoolFace
Datasetpublic

alkaid0618/ChineseWebText

ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText. ChineseWebText Dataset Overview We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is… See the full description on the dataset page: https://huggingface.co/datasets/alkaid0618/ChineseWebText.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes967downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

alkaid0618/ChineseWebText · main · files are served by the source, never re-hosted here