CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SwayStar123 /preprocessed_commoncatalog-cc-byI also seperately provide just the prompts in prompts.json keys are the image_id, and the values are the captions generated Captions generated by moondream: vikhyatk/moondream2 Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384 Original dataset: common-canvas/commoncatalog-cc-by Latents f32 and embeddings are f16 bytes Compute cost: 16x3090 for 3 day. Approximately. text10M<n<100M4 likes671k downloads2y agoHugging Face02tau /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.textquestion-answering10K<n<100K155 likes278k downloads3y agoHugging Face03fixie-ai /common_voice_17_0audio10M<n<100M18 likes216k downloads2y agoHugging Face04PleIAs /common_corpus Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.tabular10K<n<100K423 likes198k downloads5mo agoHugging Face05common-canvas /commoncatalog-cc-by Dataset Card for CommonCatalog CC-BY This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100 million… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by.imagetext-to-image10M<n<100M41 likes53k downloads2y agoHugging Face06common-canvas /commoncatalog-cc-by-nc-sa Dataset Card for CommonCatalog CC-BY-NC-SA This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc-sa.imagetext-to-image10M<n<100M5 likes47k downloads2y agoHugging Face07common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes26k downloads1y agoHugging Face08common-canvas /commoncatalog-cc-by-nc Dataset Card for CommonCatalog CC-BY-NC This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc.imagetext-to-image10M<n<100M8 likes21k downloads2y agoHugging Face09extraordinarylab /commonsense-qatext10K<n<100K0 likes20k downloads11mo agoHugging Face10SwayStar123 /preprocessed_DCAE-f64_1024_commoncatalog-cc-bytext10M<n<100M0 likes19k downloads1y agoHugging Face11common-canvas /commoncatalog-cc-by-sa Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-sa.imagetext-to-image1M<n<10M12 likes15k downloads2y agoHugging Face12moca-embed /pixelprose_commonpool Pixelprose-commonpool used in MoCa Continual Pre-training 🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper Introduction This is a interleaved multimodal pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from the commonpool split of Pixelprose by concatenating VLM captions generated by Gemini and the oringal images. The dataset consists of interleaved multimodal examples. text… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/pixelprose_commonpool.text1M<n<10M0 likes14k downloads1y agoHugging Face13commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes12k downloads10mo agoHugging Face14musabg /commoncrawl-tr Dataset Card for "commoncrawl-tr" More Information needed text10M<n<100M4 likes9.5k downloads3y agoHugging Face15wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B0 likes8.1k downloads1mo agoHugging Face16common-pile /stackv2 Stack V2 Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing this… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2.texttext-generation1M<n<10M0 likes7.9k downloads1y agoHugging Face17commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.4k downloads11d agoHugging Face18sailor2 /sea-commoncrawltext100M<n<1B1 likes7.1k downloads2y agoHugging Face19common-pile /stackv2_edu_filtered Stack V2 Edu Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.tabulartext-generation10M<n<100M6 likes6.9k downloads1y agoHugging Face20common-canvas /commoncatalog-cc-by-nd Dataset Card for CommonCatalog CC-BY-ND This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nd.imagetext-to-image1M<n<10M2 likes6.5k downloads2y agoHugging Face21agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes6.3k downloads2y agoHugging Face22common-pile /caselaw_access_project Caselaw Access Project Description This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener. The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years. In addition, Court Listener adds over 900 thousand cases scraped from 479 courts. The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project.texttext-generation1M<n<10M220 likes6.3k downloads1y agoHugging Face23common-pile /raw_v0.1_parquet Common Pile v0.1 — Parquet Consolidated Description This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection, converted to Apache Parquet and consolidated in a single repository. Nothing has been filtered or modified; the only changes are: Format: original JSON → Parquet Layout: many repositories → one consolidated dataset Extra column: a len_category bucket for quick length-based filtering Only the three original columns (id, text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/raw_v0.1_parquet.texttext-generation1B<n<10B1 likes6.2k downloads1y agoHugging Face24permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.9k downloads2y agoHugging Face25OpenVideo /Youtube-Common-First-600-Parquettextn<1K0 likes5k downloads2y agoHugging Face26common-pile /arxiv_papers ArXiv Papers Description ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline: first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.texttext-generation100K<n<1M17 likes4.8k downloads1y agoHugging Face27BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B41 likes4.6k downloads1y agoHugging Face28common-pile /peS2o_filtered PeS2o Description This dataset is a version of the peS2o dataset restricted to openly licensed articles. PeS2o is derived from S2ORC, a corpus of openly licensed abstract and full-text papers that have been converted to a structured format using Grobid. Starting from Grobid’s XML output, peS2o filters papers that are too short, have incorrect metadata, are in languages other than English, and contain OCR errors using a combination of heuristic- and model-based filtering… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/peS2o_filtered.texttext-generation1M<n<10M4 likes4.4k downloads1y agoHugging Face29common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.3k downloads1y agoHugging Face30oahegiaerhg /common_corpus Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/oahegiaerhg/common_corpus.tabular10K<n<100K0 likes4.3k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.