CoolFace
Datasetpublic

agentlans/common-crawl-sample

Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.

sourceHugging Faceupdated 2y agoView on Hugging Face
8likes5.6kdownloads
8 commits on main
0c00b202y ago

Update README.md

agentlans
75efe402y ago

Update README.md

agentlans
bdd2db22y ago

Update README.md

agentlans
955d9072y ago

Update README.md

agentlans
787c8772y ago

Update README.md

agentlans
6e3573d2y ago

Create README.md

agentlans
c617bcc2y ago

Upload 226 files

agentlans
48da58c2y ago

initial commit

agentlans