CoolFace
Datasetpublic

BramVanroy/CommonCrawl-CreativeCommons

The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur!… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.

sourceHugging Faceccupdated 1y agoView on Hugging Face
41likes4.4kdownloads
settings

This repository belongs to BramVanroy on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCommonCrawl-CreativeCommons
visibilitypublic
licencecc
gatedno
ownerBramVanroy
Account settings
BramVanroy/CommonCrawl-CreativeCommons · CoolFace