monica-sekoyan/TickTockVQA-segmented
TickTockVQA Segmented (SAM3 crops) Real world analog clock images from jaeha-choi/TickTockVQA, each cropped to the clock face by a SAM3 segmentation pass, packaged for grounded visual reasoning experiments with vision language models. Every image is paired with a single fixed instruction and a ground truth time label. All answers are in H:MM format: this corpus contains no second hand, so it exercises real world hour and minute reading only and carries no H:MM:SS signal.… See the full description on the dataset page: https://huggingface.co/datasets/monica-sekoyan/TickTockVQA-segmented.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face