CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SwayStar123 /preprocessed_commoncatalog-cc-byI also seperately provide just the prompts in prompts.json keys are the image_id, and the values are the captions generated Captions generated by moondream: vikhyatk/moondream2 Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384 Original dataset: common-canvas/commoncatalog-cc-by Latents f32 and embeddings are f16 bytes Compute cost: 16x3090 for 3 day. Approximately. text10M<n<100M4 likes671k downloads2y agoHugging Face02tau /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.textquestion-answering10K<n<100K155 likes278k downloads3y agoHugging Face03fixie-ai /common_voice_17_0audio10M<n<100M18 likes216k downloads2y agoHugging Face04PleIAs /common_corpus Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.tabular10K<n<100K423 likes198k downloads5mo agoHugging Face05common-canvas /commoncatalog-cc-by Dataset Card for CommonCatalog CC-BY This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100 million… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by.imagetext-to-image10M<n<100M41 likes53k downloads2y agoHugging Face06common-canvas /commoncatalog-cc-by-nc-sa Dataset Card for CommonCatalog CC-BY-NC-SA This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc-sa.imagetext-to-image10M<n<100M5 likes47k downloads2y agoHugging Face07common-canvas /commoncatalog-cc-by-nc Dataset Card for CommonCatalog CC-BY-NC This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc.imagetext-to-image10M<n<100M8 likes21k downloads2y agoHugging Face08extraordinarylab /commonsense-qatext10K<n<100K0 likes20k downloads11mo agoHugging Face09SwayStar123 /preprocessed_DCAE-f64_1024_commoncatalog-cc-bytext10M<n<100M0 likes19k downloads1y agoHugging Face10common-canvas /commoncatalog-cc-by-sa Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-sa.imagetext-to-image1M<n<10M12 likes15k downloads2y agoHugging Face11moca-embed /pixelprose_commonpool Pixelprose-commonpool used in MoCa Continual Pre-training 🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper Introduction This is a interleaved multimodal pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from the commonpool split of Pixelprose by concatenating VLM captions generated by Gemini and the oringal images. The dataset consists of interleaved multimodal examples. text… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/pixelprose_commonpool.text1M<n<10M0 likes14k downloads1y agoHugging Face12commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes12k downloads10mo agoHugging Face13musabg /commoncrawl-tr Dataset Card for "commoncrawl-tr" More Information needed text10M<n<100M4 likes9.5k downloads3y agoHugging Face14wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B0 likes8.1k downloads1mo agoHugging Face15commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.4k downloads11d agoHugging Face16common-canvas /commoncatalog-cc-by-nd Dataset Card for CommonCatalog CC-BY-ND This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr. The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets. Dataset Details Dataset Description We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nd.imagetext-to-image1M<n<10M2 likes6.5k downloads2y agoHugging Face17common-pile /raw_v0.1_parquet Common Pile v0.1 — Parquet Consolidated Description This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection, converted to Apache Parquet and consolidated in a single repository. Nothing has been filtered or modified; the only changes are: Format: original JSON → Parquet Layout: many repositories → one consolidated dataset Extra column: a len_category bucket for quick length-based filtering Only the three original columns (id, text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/raw_v0.1_parquet.texttext-generation1B<n<10B1 likes6.2k downloads1y agoHugging Face18permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.9k downloads2y agoHugging Face19OpenVideo /Youtube-Common-First-600-Parquettextn<1K0 likes5k downloads2y agoHugging Face20oahegiaerhg /common_corpus Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/oahegiaerhg/common_corpus.tabular10K<n<100K0 likes4.3k downloads2mo agoHugging Face21allenai /common_gen Dataset Card for "common_gen" Dataset Summary CommonGen is a constrained text generation task, associated with a benchmark dataset, to explicitly test machines for the ability of generative commonsense reasoning. Given a set of common concepts; the task is to generate a coherent sentence describing an everyday scenario using these concepts. CommonGen is challenging because it inherently requires 1) relational reasoning using background commonsense knowledge, and 2)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/common_gen.text10K<n<100K30 likes4k downloads3y agoHugging Face22coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.9k downloads8mo agoHugging Face23SpeechTest /common_voice_16_0audio100K<n<1M0 likes3.6k downloads8mo agoHugging Face24jed351 /Traditional-Chinese-Common-Crawl-Filtered Traditional Chinese C4 Dataset Summary Data obtained from 2013~2025 Common Crawl. Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant dataset contains both simplified and traditional Chinese, which could be found here. It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset. Unfortunately, I don't have enough funding to run a deduplication across… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered.text100M<n<1B26 likes2.8k downloads1y agoHugging Face25PleIAs /Medical-Commons Medical-Commons Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias. It includes three different collection: International scientific collection of 2M articles from OpenAlex. French scientific collection of XM articles, reports and PhD theses from French institutional repositories. Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.tabular1M<n<10M2 likes2.7k downloads2y agoHugging Face26hezarai /common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0 Load # Using HF Datasets from datasets import load_dataset dataset = load_dataset("hezarai/common-voice-13-fa", split="train") # Using Hezar from hezar.data import Dataset dataset = Dataset.load("hezarai/common-voice-13-fa", split="train") audioautomatic-speech-recognition10K<n<100K1 likes2.7k downloads2y agoHugging Face27jbarrow /CommonForms CommonForms: A Large, Diverse Dataset for Form Field Detection This repository hosts the CommonForms dataset, a web-scale dataset for form field detection, introduced in the paper CommonForms: A Large, Diverse Dataset for Form Field Detection. CommonForms casts the problem of form field detection as object detection: given an image of a page, predict the location and type (Text Input, Choice Button, Signature) of form fields. Key Features: Scale: Roughly 55,000 documents comprising… See the full description on the dataset page: https://huggingface.co/datasets/jbarrow/CommonForms.imageobject-detection100K<n<1M58 likes2.7k downloads10mo agoHugging Face28jed351 /Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese. The hash based cleaned dataset can be found here. Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow) text100M<n<1B0 likes2.5k downloads1y agoHugging Face29common-pile /uspto USPTO Description In the United States, patent documents are released into the public domain as government works. Patents follow a highly standardized format with distinct required sections for background, detailed description, and claims. We include parents from the US Patents and Trademark Office (USPTO) as provided by the Google Patents Public Data dataset, which includes millions of granted patents and published patent applications dating back to 1782. We processed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/uspto.texttext-generation10M<n<100M1 likes2.3k downloads1y agoHugging Face30willcai /wav2vec2_common_voice_accents_3tabular100K<n<1M0 likes2.3k downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.