CoolFace
Datasetpublic

open-index/ccrawl-urls

Common Crawl URL Index Every URL Common Crawl has seen, as a slim columnar table, ready to seed a crawler frontier What is it? This dataset is the URL-level index of Common Crawl, republished as clean Parquet. Common Crawl is a non-profit that crawls the web every month and freely publishes its archives. Each crawl ships a columnar URL index that lists every captured page with its host, fetch status, content type, detected language, and a pointer into the WARC… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-urls.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
0likes735downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
open-index/ccrawl-urls · CoolFace