CoolFace
Datasetpublic

beyond/chinese_clean_passages_80m

chinese_clean_passages_80m 包含8千余万(88328203)个纯净中文段落,不包含任何字母、数字。Containing more than 80 million pure & clean Chinese passages, without any letters/digits/special tokens. 文本长度大部分介于50~200个汉字之间。The passage length is approximately 50~200 Chinese characters. 通过datasets.load_dataset()下载数据,会产生38个大小约340M的数据包,共约12GB,所以请确保有足够空间。Downloading the dataset will result in 38 data shards each of which is about 340M and 12GB in total. Make sure there's enough space in your device:) >>>… See the full description on the dataset page: https://huggingface.co/datasets/beyond/chinese_clean_passages_80m.

sourceHugging Faceupdated 4y agoView on Hugging Face
30likes4.3kdownloads

beyond/chinese_clean_passages_80m · main · files are served by the source, never re-hosted here