CoolFace
Datasetpublic

MCG-NJU/CaReBench

CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang πŸ€— Model    |    πŸ€— Data   ο½œ    πŸ“‘ Paper    πŸ“ Introduction 🌟 CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/CaReBench.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes237downloads
Dataset Card

<div align="center"> <h1 style="margin: 0"> <img src="assets/logo.png" style="width:1.5em; vertical-align: middle; display: inline-block; margin: 0" alt="Logo"> <span style="vertical-align: middle; display: inline-block; margin: 0"><b>CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval</b></span> </h1>

<p style="margin: 0"> Yifan Xu, <a href="https://scholar.google.com/citations?user=evR3uR0AAAAJ">Xinhao Li</a>, Yichun Yang, Desen Meng, Rui Huang, <a href="https://scholar.google.com/citations?user=HEuN8PcAAAAJ">Limin Wang</a> </p>

<p align="center"> πŸ€— <a href="https://huggingface.co/MCG-NJU/CaRe-7B">Model</a> &nbsp&nbsp | &nbsp&nbsp πŸ€— <a href="https://huggingface.co/datasets/MCG-NJU/CaReBench">Data</a> &nbsp&nbsp| &nbsp&nbsp πŸ“‘ <a href="https://arxiv.org/pdf/2501.00513">Paper</a> &nbsp&nbsp </p> </div>

[image]

πŸ“ Introduction

🌟 CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for independent spatiotemporal bias evaluation. [image]

πŸ“Š ReBias and CapST Metrics are designed specifically for retrieval and captioning tasks, providing a comprehensive evaluation framework for spatiotemporal understanding in video-language models.

⚑ CaRe: A Unified Baseline for fine-grained video retrieval and captioning, achieving competitive performance through two-stage Supervised Fine-Tuning (SFT). CaRe excels in both generating detailed video descriptions and extracting robust video features. [image]

πŸš€ State-of-the-art performance on both detailed video captioning and fine-grained video retrieval. CaRe outperforms CLIP-based retrieval models and popular MLLMs in captioning tasks. [image]