datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commoncatalog-cc-by
Dataset Card for CommonCatalog CC-BY
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100 million… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by.commoncatalog-cc-by-nc-sa
Dataset Card for CommonCatalog CC-BY-NC-SA
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc-sa.commoncatalog-cc-by-nc
Dataset Card for CommonCatalog CC-BY-NC
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc.commoncatalog-cc-by-sa
Dataset Card for CommonCatalog CC-BY-SA
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-sa.commoncatalog-cc-by-nc-nd
Dataset Card for CommonCatalog CC-BY-NC-ND
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nc-nd.commoncatalog-cc-by-nd
Dataset Card for CommonCatalog CC-BY-ND
This dataset is a large collection of high-resolution Creative Common images (composed of different licenses, see paper Table 1 in the Appendix) collected in 2014 from users of Yahoo Flickr.
The dataset contains images of up to 4k resolution, making this one of the highest resolution captioned image datasets.
Dataset Details
Dataset Description
We provide captions synthetic captions to approximately 100… See the full description on the dataset page: https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by-nd.african-common-cropsCommonForms
CommonForms: A Large, Diverse Dataset for Form Field Detection
This repository hosts the CommonForms dataset, a web-scale dataset for form field detection, introduced in the paper CommonForms: A Large, Diverse Dataset for Form Field Detection.
CommonForms casts the problem of form field detection as object detection: given an image of a page, predict the location and type (Text Input, Choice Button, Signature) of form fields.
Key Features:
Scale: Roughly 55,000 documents comprising… See the full description on the dataset page: https://huggingface.co/datasets/jbarrow/CommonForms.commonforms_val_subset
Dataset Card for CommonForms_val
CommonForms_val is a validation subset of the CommonForms dataset for form field detection. It contains 10,000 annotated document images with bounding boxes for three types of form fields: text inputs, choice buttons (checkboxes/radio buttons), and signature fields. This dataset is designed for training and evaluating object detection models on the task of automatically detecting fillable form fields in document images.
This is a FiftyOne dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/commonforms_val_subset.CommonSketch
CommonSketch
Dataset Summary
CommonSketch is a semantically annotated sketch dataset introduced in the paper SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering. The dataset contains 23,100 human-drawn sketches across 300 object classes. Each sketch is paired with a fine-grained caption and element-level commonsense annotations for evaluating sketch abstraction and semantic recognizability.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ziiio/CommonSketch.wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.Common-O
Common-O
measuring multimodal reasoning across scenes
Common-O, inspired by cognitive tests for humans, probes multimodal LLMs' ability to reason across scenes by asking "what’s in common?"
Common-O is comprised of household objects:
We have two subsets: Common-O (3 - 8 objects) and Common-O Complex (8 - 16 objects).
Multimodal LLMs excel at single image perception, but struggle with multi-scene reasoning
Evaluating a Multimodal LLM on Common-O
import… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Common-O.wikimedia-commons-maps_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-maps reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-maps_beir.vel_commons_wikidata
Visual Entity Linking: Wikimedia Commons & Wikidata
This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict.
Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/
uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.chinese_fonts_common_128x128
Dataset Card for "chinese_fonts_common_128x128"
More Information needed
TTIC-commonwikimedia-commons-documents-ml_deprecated
Wikimedia Commons Document Retrieval
Wikimedia Commons Documents
This dataset is created for the evaluation of retrieval models. It contains images of (mostly historic) documents which should be identified based on their description. We extracted those descriptions from Wikimedia Commons. We have included the license type and a link (license_text) to the original Wikimedia Commons page for each extracted image.
The text_description column contains OCR text extracted from the images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_deprecated.safe-commons-pd-3m
Safe Commons PD 3M
This is a balanced and safe-to-use public domain / CC0 images dataset.
All images and texts come from Wikimedia Commons and Wikidata with strict filtering.
Images license is either Public Domain or CC0 (varies by image).
Texts license is either CC0 or CC BY-SA (varies by caption source).
No synthetic data (AI generated images or captions) is in the dataset.
To build this dataset, we tried to avoid any knowledge leaks from existing pre-trained models at the… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/safe-commons-pd-3m.CommonForms
CommonForms: A Large, Diverse Dataset for Form Field Detection
This repository hosts the CommonForms dataset, a web-scale dataset for form field detection, introduced in the paper CommonForms: A Large, Diverse Dataset for Form Field Detection.
CommonForms casts the problem of form field detection as object detection: given an image of a page, predict the location and type (Text Input, Choice Button, Signature) of form fields.
Key Features:
Scale: Roughly 55,000 documents comprising… See the full description on the dataset page: https://huggingface.co/datasets/WEwoCram/CommonForms.commonpool-128-dinov2-small
CommonPool-128-DINOv2-small
This is a derived WebDataset version of
quinnlue/commonpool-256-ssl with 128x128
stored images and precomputed facebook/dinov2-small image embeddings.
Provenance
Source dataset: quinnlue/commonpool-256-ssl at revision 78cc60b3ba29bd3fe3bd212eb22faefc4906471d
Source license/use terms: non-commercial research, inherited from the source dataset
Model: facebook/dinov2-small at revision ed25f3a31f01632728cabb09d1542f84ab7b0056
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/commonpool-128-dinov2-small.CommonObjectsBench
CommonObjectsBench: A Benchmark Dataset for General Object Image Retrieval
Dataset Description
CommonObjectsBench is a benchmark dataset for evaluating image retrieval systems on general objects and common scenes. The dataset consists of natural language queries paired with images, along with binary relevance labels indicating whether each image is relevant to the query. The dataset is designed to test retrieval systems' ability to find relevant images based on queries… See the full description on the dataset page: https://huggingface.co/datasets/sagecontinuum/CommonObjectsBench.CommonForms
CommonForms: A Large, Diverse Dataset for Form Field Detection
This repository hosts the CommonForms dataset, a web-scale dataset for form field detection, introduced in the paper CommonForms: A Large, Diverse Dataset for Form Field Detection.
CommonForms casts the problem of form field detection as object detection: given an image of a page, predict the location and type (Text Input, Choice Button, Signature) of form fields.
Key Features:
Scale: Roughly 55,000 documents comprising… See the full description on the dataset page: https://huggingface.co/datasets/kurianmelvin/CommonForms.chinese_fonts_common_512x512
Dataset Card for "chinese_fonts_common_512x512"
More Information needed
recycling-in-common-contextCommonGameCorruptionscommonsense-baselinecommonSpiderscommoncatalog-cc-by-recap-qwen3p5-35b-a3b
CommonCatalog CC-BY recaptions with Qwen3.5-35B-A3B
Dataset commoncatalog-cc-by: 14.577 Million caption rows.
This public caption-only repository contains 14,576,560 generated captions for 14,576,558 image assets and no image payload. It includes 2 additional distinct caption variants. Rows match the public image-bearing common-canvas/commoncatalog-cc-by release at revision 80f50fe4a1ca937f37a11be3f8eee5199d776ff3 through Flickr photoid, represented here by asset_instance_id… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/commoncatalog-cc-by-recap-qwen3p5-35b-a3b.gimp_common_filescommon_pool
