CoolFace
Datasetpublic

thaotien/movies_CLIP_ViT-L14

๐ŸŽฌ Movie Frame & Caption Dataset ๐Ÿ“– Introduction This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted. This dataset can be used for tasks such as: Video understanding Multimodal learning (image + text) Image captioning Vision-language retrieval ๐Ÿ“‚โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes35downloads

thaotien/movies_CLIP_ViT-L14 ยท main ยท files are served by the source, never re-hosted here