datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TripleSumm-MoSu
Dataset Summary
MoSu (Most Replayed Multimodal Video Summarization) is the first large-scale multimodal video summarization dataset. It provides synchronized visual, audio, and text features for 52,678 in-the-wild videos. The ground-truth annotations are based on YouTube's "Most Replayed" statistics, offering highly reliable per-frame importance scores derived from collective viewer engagement.
Paper: TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
GitHub… See the full description on the dataset page: https://huggingface.co/datasets/hminjeong/TripleSumm-MoSu.TripleSumm-Mr.HiSum
Dataset Summary
The original MR.HiSum (Most-replayed Highlight Detection and Summarization) was designed as a unimodal dataset and only provides pre-extracted features, which limits its use for multimodal research. To support the multimodal video summarization approach proposed in TripleSumm, we reconstructed the dataset by independently crawling the original videos using the provided metadata and extracting features across three distinct modalities: Visual, Audio, and Text.
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/hminjeong/TripleSumm-Mr.HiSum.TacVLBench
