datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pixmo-docs
PixMo-Docs
We now recommend using CoSyn-400k and CoSyn-point over these
datasets. They are improved versions with more images categories and an improved generation pipeline.
PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.pixmo-cap-images
PixMo-Cap
Big thanks to Ai2 for releasing the original PixMo-Cap dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images.
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models.
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model… See the full description on the dataset page: https://huggingface.co/datasets/anthracite-org/pixmo-cap-images.pixmo_images
PixMo Images
The raw images backing the PixMo
datasets used to train Molmo, packaged
as Parquet shards with embedded image bytes so they can be browsed in the dataset viewer
and loaded directly with datasets.
The PixMo annotation datasets (allenai/pixmo-*) ship image_urls rather than image
bytes. This repository is a content cache of those images, keyed by the SHA-256 of the
source URL.
Contents
1,073,189 images across 525 Parquet shards (data/train-*.parquet)… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/pixmo_images.pixmo-cap
PixMo-Cap
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models.
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model to turn the audio transcripts(s) into a long caption.
The audio transcripts are also included.
PixMo-Cap is part of the PixMo dataset collection and was used to train the Molmo family of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-cap.pixmo-points
PixMo-Points
PixMo-Points is a dataset of images paired with referring expressions and points marking the locations the
referring expression refers to in the image. It was collected using human annotators and contains a diverse
range of points and expressions, with many high-frequency (10+) expressions.
PixMo-Points is a part of the PixMo dataset collection and was used to
provide the pointing capabilities of the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-points.pixmo-count
PixMo-Count
PixMo-Count is a dataset of images paired with objects and their point locations in the image.
It was built by running the Detic object detector on web images, and then filtering the data
to improve accuracy and diversity. The val and test sets are human-verified and only contain counts from 2 to 10.
PixMo-Count is a part of the PixMo dataset collection and was used to
augment the pointing capabilities of the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-count.pixmo-cap-imagespixmo-pointsar_pixmocapqatrans_instructpixmo-cap-qa-imagespixmo-point-count-concat_0-20qwen3-resize-easyr1-110k-bbox0p05-remove-pixmo-uground-seeclickpixmo-point-count-gen-undar_pixmocapqatrans2_instructpixmo-cap-qa-imagesBig thanks to Ai2 for releasing the original PixMo-CapQA dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images.
PixMo-CapQA
PixMo-CapQA is a synthetic dataset of question/answer pairs about images. The data was generated by using the
Claude large language model to build Q/A pairs from dense captions of images (the model did not see the actual images).
PixMo-CapQA is a part of the PixMo dataset collection… See the full description on the dataset page: https://huggingface.co/datasets/anthracite-org/pixmo-cap-qa-images.pixmo-ask-model-anything
PixMo-AskModelAnything
PixMo-AskModelAnything is an instruction-tuning dataset for vision-language models. It contains human-authored
question-answer pairs about diverse images with long-form answers.
PixMo-AskModelAnything is a part of the PixMo dataset collection and was used to train the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos
Loading
data = datasets.load_dataset("allenai/pixmo-ask-model-anything", split="train")
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-ask-model-anything.pixmo-cap-images
PixMo-Cap
Big thanks to Ai2 for releasing the original PixMo-Cap dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images.
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models.
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/pixmo-cap-images.ro_sft_pixmo_cap
Dataset Description
PixmoCap is a dataset of very long (roughly 200 words on average), detailed captions.
Here we provide the Romanian translation of the PixmoCap dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@inproceedings{deitke2025molmo,
title={Molmo and pixmo: Open weights and… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_cap.ro_sft_pixmo_points
Dataset Description
PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image.
Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.pixmo-ask-model-anything-imagespixmo-cap-qa
PixMo-CapQA
PixMo-CapQA is a synthetic dataset of question/answer pairs about images. The data was generated by using the
Claude large language model to build Q/A pairs from dense captions of images (the model did not see the actual images).
PixMo-CapQA is a part of the PixMo dataset collection and was used to train the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos
Loading
data = datasets.load_dataset("allenai/pixmo-cap-qa", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-cap-qa.pixmo-cap-imagesar_pixmodocsother_instructpixmo-points-filtered-below10_imgContainedpixmopoints-imagespixmo-points-evalro_sft_pixmo_aa
Dataset Description
PixmoAA is an instruction-tuning dataset for vision-language models. It contains human-authored question-answer pairs about diverse images with long-form answers.
Here we provide the Romanian translation of the PixmoAA dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_aa.pixmo-count-filtered-imgContainedpixmo-points-eval
PixMo-Points-Eval
PixMo-Points-Eval is a subset of PixMo-Points that has been human-filtered and annotated with segmentation masks.
It is used for pointing evaluations.
PixMo-Points is a part of the PixMo dataset collection and was used to
provide the pointing capabilities of the Molmo family of models
Loading
data = datasets.load_dataset("allenai/pixmo-points-eval", split="test")
Data Format
Images are stored as URLs that will need to be downloaded… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-points-eval.easyr1-103k-4MP-not-all-correct-stage-one-temp-1_1-RL-remove-pixmo-uground-seeclick-refusal
