prithivMLmods/VLM-Video-Understanding
VLM-Video-Understanding A minimalistic demo for image inference and video understanding using OpenCV, built on top of several popular open-source Vision-Language Models (VLMs). This repository provides Colab notebooks demonstrating how to apply these VLMs to video and image tasks using Python and Gradio. Overview This project showcases lightweight inference pipelines for the following: Video frame extraction and preprocessing Image-level inference with VLMs… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/VLM-Video-Understanding.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face