mihaimasala/human_rankings
This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper. Dataset Description The dataset comprises human preference rankings of automatically generated video descriptions. For each video, users were asked to rank five descriptions based on criteria such as completeness and richness. Citation @misc{masala2025visionlanguagegraphevents, title={From Vision To Language through… See the full description on the dataset page: https://huggingface.co/datasets/mihaimasala/human_rankings.
This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper.
Dataset Description
The dataset comprises human preference rankings of automatically generated video descriptions. For each video, users were asked to rank five descriptions based on criteria such as completeness and richness.
Citation
@misc{masala2025visionlanguagegraphevents,
title={From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach},
author={Mihai Masala and Marius Leordeanu},
year={2025},
eprint={2507.04815},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.04815},
}