CoolFace
Datasetpublic

monology/c5-en-filtered

This is the 2022-05 snapshot of BramVanroy/CommonCrawl-CreativeCommons, filtered by: Extracting the URLs from the dataset Getting documents that match those URLs from the corresponding snapshot of togethercomputer/RedPajama-Data-V2 Keeping only the head and middle partitions of ccnet Keeping documents with at least 50 words and a mean word length between 3 and 10 inclusive In total we keep 4,553,263 of the 15,239,155 total documents.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes49downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
monology/c5-en-filtered · CoolFace