datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video_captioningMSVD-Video-Captioning-Vi
MSVD-Video-Captioning-Vi
📌 Overview
MSVD-Video-Captioning-Vi is a Vietnamese video captioning dataset derived from the MSVD dataset originally hosted by the user friedrichor on Hugging Face.
This dataset provides Vietnamese captions for short video clips and is intended for:
Video captioning research
Vision–Language model training
Multimodal instruction tuning
Video-to-text generation
🔁 Dataset Origin
This dataset is a translated and derived version… See the full description on the dataset page: https://huggingface.co/datasets/NTQAI/MSVD-Video-Captioning-Vi.Video-Captioning-Generation-Benchmark-ExampleMSR-VTT-Video-Captioning-Vivideo-captioning
