monica-sekoyan/TickTockVQA-segmented
TickTockVQA Segmented (SAM3 crops) Real world analog clock images from jaeha-choi/TickTockVQA, each cropped to the clock face by a SAM3 segmentation pass, packaged for grounded visual reasoning experiments with vision language models. Every image is paired with a single fixed instruction and a ground truth time label. All answers are in H:MM format: this corpus contains no second hand, so it exercises real world hour and minute reading only and carries no H:MM:SS signal.… See the full description on the dataset page: https://huggingface.co/datasets/monica-sekoyan/TickTockVQA-segmented.
0197
