creative-commons
CommonCrawl-CreativeCommons
The Common Crawl Creative Commons Corpus (C5)
Raw CommonCrawl crawls, annotated with Creative Commons license information
C5 is an effort to collect Creative Commons-licensed web data in one place.
The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.CommonCrawl-CreativeCommons-fine
Common Crawl Creative Commons Corpus Fine (C5f)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5.
Created with this script.
For more information, see C5.
Progress
In the v1 release, the following crawls are included
CC-MAIN-2019-30
CC-MAIN-2020-05CC-MAIN-2023-06
CC-MAIN-2024-51
CC-MAIN-2024-46
CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.CreativeCommons-RAG-QA-Mixtral8x22b
以下のデータ源からランダムに抽出した日本語のテキストをもとに、RAG形式のQ&Aを自動生成したものです。
Wikibooks
Wikipedia
判例データ
instruction datasetとしてではなく、事前学習での利用を想定しています(質疑応答をするための訓練)。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
