CoolFace
4 results

creative-commons

BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B41 likes4.4k downloads1y agoHugging FaceBramVanroy /CommonCrawl-CreativeCommons-fine Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46 CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.texttext-generation10M<n<100M5 likes996 downloads1y agoHugging FaceBramVanroy /CommonCrawl-CreativeCommons-strict Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.texttext-generation10M<n<100M2 likes478 downloads1y agoHugging Facekanhatakeyama /CreativeCommons-RAG-QA-Mixtral8x22b 以下のデータ源からランダムに抽出した日本語のテキストをもとに、RAG形式のQ&Aを自動生成したものです。 Wikibooks Wikipedia 判例データ instruction datasetとしてではなく、事前学習での利用を想定しています(質疑応答をするための訓練)。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text1M<n<10M0 likes12 downloads2y agoHugging Face