agentlans/common-crawl-sample
Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.
85.6k
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md
Upload 226 files
initial commit
