datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peek_vqa
Peek VQA Dataset Card
Dataset details
This dataset contains 2M image-QA pairs used to fine-tune PEEK VLM, a VLM for robotics that answers "where policies should focus" and "what policies should do".
Given an instruction (<quest></quest>), the task is to predict a unified point-based representation corresponding to 1) a path guiding the robot end-effector in what actions to take (TRAJECTORY), and 2) a set of task-relevant masking points that show where to focus on (MASK).… See the full description on the dataset page: https://huggingface.co/datasets/memmelma/peek_vqa.ocr-vqa-200k_imagesImage collections for OCR-VQA-200K.
Image size: 208,467.
OCR_VQAKADIS700KST-VQAKADISvqamscoco-vqattack-eps-sweep
