CoolFace
Datasetpublic

mihaimasala/human_rankings

This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper. Dataset Description The dataset comprises human preference rankings of automatically generated video descriptions. For each video, users were asked to rank five descriptions based on criteria such as completeness and richness. Citation @misc{masala2025visionlanguagegraphevents, title={From Vision To Language through… See the full description on the dataset page: https://huggingface.co/datasets/mihaimasala/human_rankings.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes29downloads
Dataset Card

This dataset supports the From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach paper.

Dataset Description

The dataset comprises human preference rankings of automatically generated video descriptions. For each video, users were asked to rank five descriptions based on criteria such as completeness and richness.

Citation

@misc{masala2025visionlanguagegraphevents,
      title={From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach}, 
      author={Mihai Masala and Marius Leordeanu},
      year={2025},
      eprint={2507.04815},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2507.04815}, 
}