CoolFace
18 results

gqa

lmms-lab-encoder /GQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{hudson2019gqa, title={Gqa: A new dataset for real-world visual reasoning and compositional… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/GQA.image10M<n<100M34 likes39k downloads3y agoHugging FaceVoxel51 /GQA-Scene-Graph Dataset Card for GQA-35k The GQA (Visual Reasoning in the Real World) dataset is a large-scale visual question answering dataset that includes scene graph annotations for each image. This is a FiftyOne dataset with 35000 samples. Note: This is a 35,000 sample subset which does not contain questions, only the scene graph annotations as detection-level attributes. You can find the recipe notebook for creating the dataset here Installation If you haven't already… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/GQA-Scene-Graph.imageobject-detection10K<n<100K4 likes2.5k downloads2y agoHugging Facejinyoungkim /NExT-GQA Can I Trust Your Answer? Visually Grounded Video Question Answering Introduction We study visually grounded VideoQA by forcing vision-language models (VLMs) to answer questions and simultaneously ground the relevant video moments as visual evidences. We show that this task is easy for human yet is extremely challenging for existing VLMs, revealing that the strong QA performance of these models may largely due to short-cut learning (e.g., language priors and spurious vision-text… See the full description on the dataset page: https://huggingface.co/datasets/jinyoungkim/NExT-GQA.2 likes2.4k downloads1y agoHugging Facepppop7 /GQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{hudson2019gqa, title={Gqa: A new dataset for real-world visual reasoning and compositional question… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/GQA.image10M<n<100M0 likes913 downloads9mo agoHugging Facevikhyatk /gqaimage10K<n<100K2 likes585 downloads2y agoHugging Facegligen /gqa_tsv0 likes498 downloads3y agoHugging Face