molmo2
Datasets
All datasets matching “molmo2”Molmo2-SynMultiImageQA
Molmo2-SynMultiImageQA
Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images, including charts, tables, documents, diagrams, etc.
The synthetic data is generated by extending the CoSyn framework into multi-image settings,
with Claude-sonnet-4-5 as the coding LLM to generate code that can be executed to render an image.
Then, we use GPT-5 to generate question-answer pairs with code (without using the rendered… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-SynMultiImageQA.Molmo2-ER-VST-P
Molmo2-ER · rayruiyang/vst_500k
500K perception QA over images normalized to a uniform virtual camera (single + multi-view).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
⚠️ This dataset is released for non-commercial research use only, inheriting the most-restrictive license among its upstream sources. See the upstream repository for details.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-VST-P.Molmo2-ER-SenseNova-SI
Molmo2-ER · sensenova/SenseNova-SI-800K
832K multi-image spatial-intelligence conversations grounded in 3D scene annotations.
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: sensenova/SenseNova-SI-800K
Paper: Scaling Spatial Intelligence with Multimodal Foundation Models (arXiv:2511.13719)
License: apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-SenseNova-SI.molmo2-tulu4-classifiedMolmo2-ER-RoboPoint
Molmo2-ER · wentao-yuan/robopoint-data
1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: wentao-yuan/robopoint-data
Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721)
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.Molmo2-ER-VSI-590K
Molmo2-ER · nyu-visionx/VSI-590K
590K spatial QA samples (image+video) propagated from 3D ground truth and CV pseudo-labels.
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: nyu-visionx/VSI-590K
Paper: Cambrian-S: Towards Spatial Supersensing in Video (arXiv:2511.04670)
License: apache-2.0 (inherits from upstream)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-VSI-590K.
