video-captioning
video_captioningMulti-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.MSVD-Video-Captioning-Vi
MSVD-Video-Captioning-Vi
📌 Overview
MSVD-Video-Captioning-Vi is a Vietnamese video captioning dataset derived from the MSVD dataset originally hosted by the user friedrichor on Hugging Face.
This dataset provides Vietnamese captions for short video clips and is intended for:
Video captioning research
Vision–Language model training
Multimodal instruction tuning
Video-to-text generation
🔁 Dataset Origin
This dataset is a translated and derived version… See the full description on the dataset page: https://huggingface.co/datasets/NTQAI/MSVD-Video-Captioning-Vi.valorant-video-captioning-1.1-dMSR-VTT-Video-Captioning-ViVideo-Captioning-Generation-Benchmark-Example
