CoolFace
Datasetpublic

mihaimasala/vlm_jury_rankings

This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper. Dataset Description The dataset comprises VLMs preference rankings of automatically generated video descriptions. For each video, the VLMs were asked to rank six descriptions based on criteria such as completeness and richness. Citation @misc{masala2025visionlanguagegraphevents, title={From Vision To Language through… See the full description on the dataset page: https://huggingface.co/datasets/mihaimasala/vlm_jury_rankings.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes22downloads
Dataset Card

This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper.

Dataset Description

The dataset comprises VLMs preference rankings of automatically generated video descriptions. For each video, the VLMs were asked to rank six descriptions based on criteria such as completeness and richness.

Citation

@misc{masala2025visionlanguagegraphevents,
      title={From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach}, 
      author={Mihai Masala and Marius Leordeanu},
      year={2025},
      eprint={2507.04815},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2507.04815}, 
}