CoolFace
Datasetpublic

Symato/cc

What is Symato CC? To download all WARC data from Common Crawl then filter out Vietnamese in Markdown and Plaintext format. There is 1% of Vietnamse in CC, extract all of them out should be a lot (~10TB of plaintext). Main contributors https://huggingface.co/nampdn-ai https://huggingface.co/binhvq https://huggingface.co/th1nhng0 https://huggingface.co/iambestfeed Simple quality filters To make use of raw data from common crawl, you need to do… See the full description on the dataset page: https://huggingface.co/datasets/Symato/cc.

sourceHugging Facemitupdated 3y agoView on Hugging Face
3likes414kdownloads
settings

This repository belongs to Symato on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namecc
visibilitypublic
licencemit
gatedno
ownerSymato
Account settings
Symato/cc · CoolFace