prithivMLmods/VLM-Video-Understanding
VLM-Video-Understanding A minimalistic demo for image inference and video understanding using OpenCV, built on top of several popular open-source Vision-Language Models (VLMs). This repository provides Colab notebooks demonstrating how to apply these VLMs to video and image tasks using Python and Gradio. Overview This project showcases lightweight inference pipelines for the following: Video frame extraction and preprocessing Image-level inference with VLMs… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/VLM-Video-Understanding.
292
