microsoft/IMAGE_UNDERSTANDING
A key question for understanding multimodal performance is analyzing the ability for a model to have basic vs. detailed understanding of images. These capabilities are needed for models to be used in real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting. The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/IMAGE_UNDERSTANDING.
Update README.md
Update README.md
Update README.md
Upload 8 files
Update README.md
Update README.md
Update README.md
Update README.md
Upload 8 files
Upload visual_prompting_val.parquet
Upload visual_prompting_val.parquet
Upload 2 files
Upload 2 files
Delete spatial_reasoning_lrtb_pairs/msr_aif_spatial_reasoning_lrtb_pairs.parquet
Upload 2 files
Upload object_detection_val_long_prompt.parquet
Upload object_detection_val_long_prompt.parquet
Update README.md
Update README.md
Update README.md
Upload coco_instances.json
Upload coco_instances.json
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload 8 files
Update README.md
Delete msr_aif_visual_prompting_pairs
Delete msr_aif_visual_prompting_single
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload visual_prompting_val.parquet
Update README.md
