datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
preprocessed_commoncatalog-cc-byI also seperately provide just the prompts in prompts.json
keys are the image_id, and the values are the captions generated
Captions generated by moondream: vikhyatk/moondream2
Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix
Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384
Original dataset: common-canvas/commoncatalog-cc-by
Latents f32 and embeddings are f16 bytes
Compute cost: 16x3090 for 3 day. Approximately.
commonsense_qa
Dataset Card for "commonsense_qa"
Dataset Summary
CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge
to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers.
The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation
split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.common_voice_17_0common_corpus
Common Corpus
Full paper - ICLR 2026 oral
Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.
Common Corpus differs from existing open datasets in that it is:
Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.commoncatalog-cc-by
Dataset Card for CommonCatalog CC-BY
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100 million… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by.commoncatalog-cc-by-nc-sa
Dataset Card for CommonCatalog CC-BY-NC-SA
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc-sa.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.commoncatalog-cc-by-nc
Dataset Card for CommonCatalog CC-BY-NC
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc.commonsense-qapreprocessed_DCAE-f64_1024_commoncatalog-cc-bycommoncatalog-cc-by-sa
Dataset Card for CommonCatalog CC-BY-SA
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-sa.pixelprose_commonpool
Pixelprose-commonpool used in MoCa Continual Pre-training
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
Introduction
This is a interleaved multimodal pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from the commonpool split of
Pixelprose by concatenating VLM captions generated by Gemini and the oringal images.
The dataset consists of interleaved multimodal examples. text… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/pixelprose_commonpool.gneissweb-annotation-url-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.commoncrawl-tr
Dataset Card for "commoncrawl-tr"
More Information needed
thai-commoncrawl-index
Thai Common Crawl Index (2019–2026)
An index of every page Common Crawl detected as Thai across 70 monthly crawls, from
January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30).
932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains
Each row records where the page lives inside Common Crawl's WARC archives — file name,
byte offset, and record length — so you can fetch exactly the pages you want with HTTP
range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.stackv2
Stack V2
Description
We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2.
When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list.
Per-document license information is available in the license entry of the metadata field of each example.
Code for collecting, processing, and preparing this… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2.host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.sea-commoncrawlstackv2_edu_filtered
Stack V2 Edu
Description
We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.commoncatalog-cc-by-nd
Dataset Card for CommonCatalog CC-BY-ND
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nd.common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.caselaw_access_project
Caselaw Access Project
Description
This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener.
The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years.
In addition, Court Listener adds over 900 thousand cases scraped from 479 courts.
The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project.raw_v0.1_parquet
Common Pile v0.1 — Parquet Consolidated
Description
This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection, converted to Apache Parquet and consolidated in a single repository.
Nothing has been filtered or modified; the only changes are:
Format: original JSON → Parquet
Layout: many repositories → one consolidated dataset
Extra column: a len_category bucket for quick length-based filtering
Only the three original columns (id, text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/raw_v0.1_parquet.wdc-common-crawl-embedded-jsonldYoutube-Common-First-600-Parquetarxiv_papers
ArXiv Papers
Description
ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more.
When uploading papers, authors can choose from a variety of licenses.
This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline:
first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.CommonCrawl-CreativeCommons
The Common Crawl Creative Commons Corpus (C5)
Raw CommonCrawl crawls, annotated with Creative Commons license information
C5 is an effort to collect Creative Commons-licensed web data in one place.
The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.peS2o_filtered
PeS2o
Description
This dataset is a version of the peS2o dataset restricted to openly licensed articles.
PeS2o is derived from S2ORC, a corpus of openly licensed abstract and full-text papers that have been converted to a structured format using Grobid.
Starting from Grobid’s XML output, peS2o filters papers that are too short, have incorrect metadata, are in languages other than English, and contain OCR errors using a combination of heuristic- and model-based filtering… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/peS2o_filtered.github_archive
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.common_corpus
Common Corpus
Full paper - ICLR 2026 oral
Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.
Common Corpus differs from existing open datasets in that it is:
Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/oahegiaerhg/common_corpus.
