hminjeong/TripleSumm-MoSu
Dataset Summary MoSu (Most Replayed Multimodal Video Summarization) is the first large-scale multimodal video summarization dataset. It provides synchronized visual, audio, and text features for 52,678 in-the-wild videos. The ground-truth annotations are based on YouTube's "Most Replayed" statistics, offering highly reliable per-frame importance scores derived from collective viewer engagement. Paper: TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization GitHub… See the full description on the dataset page: https://huggingface.co/datasets/hminjeong/TripleSumm-MoSu.
This repository belongs to hminjeong on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
