CoolFace
20 results

molmo2

allenai /Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images, including charts, tables, documents, diagrams, etc. The synthetic data is generated by extending the CoSyn framework into multi-image settings, with Claude-sonnet-4-5 as the coding LLM to generate code that can be executed to render an image. Then, we use GPT-5 to generate question-answer pairs with code (without using the rendered… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-SynMultiImageQA.imagevisual-question-answering100K<n<1M10 likes8.1k downloads9mo agoHugging Faceallenai /Molmo2-ER-VST-P Molmo2-ER · rayruiyang/vst_500k 500K perception QA over images normalized to a uniform virtual camera (single + multi-view). This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified. ⚠️ This dataset is released for non-commercial research use only, inheriting the most-restrictive license among its upstream sources. See the upstream repository for details.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-VST-P.text100K<n<1M0 likes4k downloads5mo agoHugging Faceallenai /Molmo2-ER-SenseNova-SI Molmo2-ER · sensenova/SenseNova-SI-800K 832K multi-image spatial-intelligence conversations grounded in 3D scene annotations. This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified. Upstream source Original dataset: sensenova/SenseNova-SI-800K Paper: Scaling Spatial Intelligence with Multimodal Foundation Models (arXiv:2511.13719) License: apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-SenseNova-SI.0 likes2.1k downloads5mo agoHugging Faceallenai /molmo2-tulu4-classifiedtext1M<n<10M1 likes1.5k downloads8mo agoHugging Faceallenai /Molmo2-ER-RoboPoint Molmo2-ER · wentao-yuan/robopoint-data 1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA). This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified. Upstream source Original dataset: wentao-yuan/robopoint-data Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721) License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.image1M<n<10M1 likes1.5k downloads5mo agoHugging Faceallenai /Molmo2-ER-VSI-590K Molmo2-ER · nyu-visionx/VSI-590K 590K spatial QA samples (image+video) propagated from 3D ground truth and CV pseudo-labels. This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified. Upstream source Original dataset: nyu-visionx/VSI-590K Paper: Cambrian-S: Towards Spatial Supersensing in Video (arXiv:2511.04670) License: apache-2.0 (inherits from upstream)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-VSI-590K.0 likes942 downloads5mo agoHugging Face