thaotien/movies_CLIP_ViT-L14
π¬ Movie Frame & Caption Dataset π Introduction This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted. This dataset can be used for tasks such as: Video understanding Multimodal learning (image + text) Image captioning Vision-language retrieval πβ¦ See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone elseβs repository from here would need an authorised integration and the account holderβs consent, so the link goes to the source instead.
Open discussions on Hugging Face