CoolFace
Datasetpublic

hthhththt/VLM-Video-Understanding

VLM-Video-Understanding A minimalistic demo for image inference and video understanding using OpenCV, built on top of several popular open-source Vision-Language Models (VLMs). This repository provides Colab notebooks demonstrating how to apply these VLMs to video and image tasks using Python and Gradio. Overview This project showcases lightweight inference pipelines for the following: Video frame extraction and preprocessing Image-level inference with VLMs… See the full description on the dataset page: https://huggingface.co/datasets/hthhththt/VLM-Video-Understanding.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes54downloads
1 commits on main
66f17246mo ago

Duplicate from prithivMLmods/VLM-Video-Understanding

hthhththt, prithivMLmods