CoolFace
Datasetpublic

alkaid0618/ChineseWebText

ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText. ChineseWebText Dataset Overview We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is… See the full description on the dataset page: https://huggingface.co/datasets/alkaid0618/ChineseWebText.

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes967downloads

No commit history came back for main. The revision may not exist, or the source declined the request.