datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GQAOur generated videos are sourced from the open-source videos of VBench, using the following link: Videos. Please use the name field in the .csv file to index the corresponding video.
The questions and options in the CSV file are numbered sequentially. For the specific questions or options corresponding to each number, please refer to index in the JSON file.
In the Original_csv folder, please change the suffix .csv_bac to .csv for use.
GQA-Q2Q
GQA-Q2Q Dataset
GQA-Q2Q is a visual question disambiguation resource built on top of the GQA dataset. The dataset focuses on scenarios where a Vision-Language Model (VLM) must identify the correct target entity among multiple visually similar or same-name entities in an image to resolve ambiguous questions.
For code to train/evaluate models and reproduce the experiments, visit the GQA-Q2Q GitHub Repository.
Quick Start
You can load this dataset directly using the… See the full description on the dataset page: https://huggingface.co/datasets/gminipark/GQA-Q2Q.
