CoolFace
Datasetpublic

BramVanroy/CommonCrawl-CreativeCommons

The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur!… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.

sourceHugging Faceccupdated 1y agoView on Hugging Face
41likes4.4kdownloads

BramVanroy/CommonCrawl-CreativeCommons · main · files are served by the source, never re-hosted here