HiTZ/MATE
Dataset Card for HiTZ/MATE This dataset provides a benchmark consisting of 5,500 question-answering examples to assess the cross-modal entity linking capabilities of vision-language models (VLMs). The ability to link entities in different modalities is measured in a question-answering setting, where each scene is represented in both the visual modality (image) and the textual one (a list of objects and their attributes in JSON format). The dataset is provided in two… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MATE.
This repository belongs to HiTZ on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
