CoolFace
Datasetpublic

darknoon/simple-shapes-svg

The goal of this dataset is to measure and improve the ability of VLMs to see accurately in spatial dimensions. I've tried to ensure that all of the examples are not too hard have sufficient contrast between foreground and background shapes are not clipped or ambiguous solid background canvas is square 512x512 Initially, I've kept the "canvas" that they're working with 512x512 points, but you can learn more by experimenting with the dimensions as well.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes67downloads
Dataset Card

The goal of this dataset is to measure and improve the ability of VLMs to see accurately in spatial dimensions.

I've tried to ensure that all of the examples are not too hard

  • —have sufficient contrast between foreground and background
  • —shapes are not clipped or ambiguous
  • —solid background
  • —canvas is square 512x512

Initially, I've kept the "canvas" that they're working with 512x512 points, but you can learn more by experimenting with the dimensions as well.