CoolFace
Datasetpublic

microsoft/IMAGE_UNDERSTANDING

A key question for understanding multimodal performance is analyzing the ability for a model to have basic vs. detailed understanding of images. These capabilities are needed for models to be used in real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting. The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/IMAGE_UNDERSTANDING.

sourceHugging Facecdla-permissive-2.0updated 2y agoView on Hugging Face
7likes2.9kdownloads
settings

This repository belongs to microsoft on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameIMAGE_UNDERSTANDING
visibilitypublic
licencecdla-permissive-2.0
gatedno
ownermicrosoft
Account settings
microsoft/IMAGE_UNDERSTANDING · CoolFace