open-index/ccrawl-urls
Common Crawl URL Index Every URL Common Crawl has seen, as a slim columnar table, ready to seed a crawler frontier What is it? This dataset is the URL-level index of Common Crawl, republished as clean Parquet. Common Crawl is a non-profit that crawls the web every month and freely publishes its archives. Each crawl ships a columnar URL index that lists every captured page with its host, fetch status, content type, detected language, and a pointer into the WARC… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-urls.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face