CoolFace
19 results

pixmo

allenai /pixmo-docs PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images categories and an improved generation pipeline. PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents. The data was created by using the Claude large language model to generate code that can be executed to render an image, and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.imagevisual-question-answering100K<n<1M35 likes3.2k downloads2y agoHugging Faceanthracite-org /pixmo-cap-images PixMo-Cap Big thanks to Ai2 for releasing the original PixMo-Cap dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images. PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions. It can be used to pre-train and fine-tune vision-language models. PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model… See the full description on the dataset page: https://huggingface.co/datasets/anthracite-org/pixmo-cap-images.imageimage-to-text100K<n<1M13 likes1.1k downloads2y agoHugging FaceUWGZQ /pixmo_images PixMo Images The raw images backing the PixMo datasets used to train Molmo, packaged as Parquet shards with embedded image bytes so they can be browsed in the dataset viewer and loaded directly with datasets. The PixMo annotation datasets (allenai/pixmo-*) ship image_urls rather than image bytes. This repository is a content cache of those images, keyed by the SHA-256 of the source URL. Contents 1,073,189 images across 525 Parquet shards (data/train-*.parquet)… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/pixmo_images.imageimage-to-text1M<n<10M1 likes985 downloads3mo agoHugging Faceallenai /pixmo-cap PixMo-Cap PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions. It can be used to pre-train and fine-tune vision-language models. PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model to turn the audio transcripts(s) into a long caption. The audio transcripts are also included. PixMo-Cap is part of the PixMo dataset collection and was used to train the Molmo family of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-cap.imageimage-to-text100K<n<1M48 likes964 downloads2y agoHugging Faceallenai /pixmo-points PixMo-Points PixMo-Points is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image. It was collected using human annotators and contains a diverse range of points and expressions, with many high-frequency (10+) expressions. PixMo-Points is a part of the PixMo dataset collection and was used to provide the pointing capabilities of the Molmo family of models Quick links: 📃 Paper 🎥 Blog with Videos… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-points.image1M<n<10M48 likes935 downloads2y agoHugging Faceallenai /pixmo-count PixMo-Count PixMo-Count is a dataset of images paired with objects and their point locations in the image. It was built by running the Detic object detector on web images, and then filtering the data to improve accuracy and diversity. The val and test sets are human-verified and only contain counts from 2 to 10. PixMo-Count is a part of the PixMo dataset collection and was used to augment the pointing capabilities of the Molmo family of models Quick links: 📃 Paper 🎥 Blog with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-count.imagevisual-question-answering10K<n<100K12 likes868 downloads2y agoHugging Face