thaotien/movies_CLIP_ViT-L14
๐ฌ Movie Frame & Caption Dataset ๐ Introduction This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted. This dataset can be used for tasks such as: Video understanding Multimodal learning (image + text) Image captioning Vision-language retrieval ๐โฆ See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.
035
