MCG-NJU/CaReBench
CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang π€ Model | π€ Data ο½ π Paper π Introduction π CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions forβ¦ See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/CaReBench.
<div align="center"> <h1 style="margin: 0"> <img src="assets/logo.png" style="width:1.5em; vertical-align: middle; display: inline-block; margin: 0" alt="Logo"> <span style="vertical-align: middle; display: inline-block; margin: 0"><b>CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval</b></span> </h1>
<p style="margin: 0"> Yifan Xu, <a href="https://scholar.google.com/citations?user=evR3uR0AAAAJ">Xinhao Li</a>, Yichun Yang, Desen Meng, Rui Huang, <a href="https://scholar.google.com/citations?user=HEuN8PcAAAAJ">Limin Wang</a> </p>
<p align="center"> π€ <a href="https://huggingface.co/MCG-NJU/CaRe-7B">Model</a>    |    π€ <a href="https://huggingface.co/datasets/MCG-NJU/CaReBench">Data</a>   ο½    π <a href="https://arxiv.org/pdf/2501.00513">Paper</a>    </p> </div>
π Introduction
π CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for independent spatiotemporal bias evaluation.
π ReBias and CapST Metrics are designed specifically for retrieval and captioning tasks, providing a comprehensive evaluation framework for spatiotemporal understanding in video-language models.
β‘ CaRe: A Unified Baseline for fine-grained video retrieval and captioning, achieving competitive performance through two-stage Supervised Fine-Tuning (SFT). CaRe excels in both generating detailed video descriptions and extracting robust video features.
π State-of-the-art performance on both detailed video captioning and fine-grained video retrieval. CaRe outperforms CLIP-based retrieval models and popular MLLMs in captioning tasks.
