mihaimasala/vlm_jury_rankings
This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper. Dataset Description The dataset comprises VLMs preference rankings of automatically generated video descriptions. For each video, the VLMs were asked to rank six descriptions based on criteria such as completeness and richness. Citation @misc{masala2025visionlanguagegraphevents, title={From Vision To Language through… See the full description on the dataset page: https://huggingface.co/datasets/mihaimasala/vlm_jury_rankings.
022
