CoolFace
Datasetpublic

monology/c5-en-filtered

This is the 2022-05 snapshot of BramVanroy/CommonCrawl-CreativeCommons, filtered by: Extracting the URLs from the dataset Getting documents that match those URLs from the corresponding snapshot of togethercomputer/RedPajama-Data-V2 Keeping only the head and middle partitions of ccnet Keeping documents with at least 50 words and a mean word length between 3 and 10 inclusive In total we keep 4,553,263 of the 15,239,155 total documents.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes49downloads
3 commits on main
0c3b7c51y ago

Create README.md

monology
e616b021y ago

Add 2022-05 snapshot

monology
7a65e331y ago

initial commit

monology