CoolFace
Datasetpublic

beyond/chinese_clean_passages_80m

chinese_clean_passages_80m 包含8千余万(88328203)个纯净中文段落,不包含任何字母、数字。Containing more than 80 million pure & clean Chinese passages, without any letters/digits/special tokens. 文本长度大部分介于50~200个汉字之间。The passage length is approximately 50~200 Chinese characters. 通过datasets.load_dataset()下载数据,会产生38个大小约340M的数据包,共约12GB,所以请确保有足够空间。Downloading the dataset will result in 38 data shards each of which is about 340M and 12GB in total. Make sure there's enough space in your device:) >>>… See the full description on the dataset page: https://huggingface.co/datasets/beyond/chinese_clean_passages_80m.

sourceHugging Faceupdated 4y agoView on Hugging Face
30likes4.3kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
beyond/chinese_clean_passages_80m · CoolFace