CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Dense-World /Sa2VA-TrainingThis repository contains the code and data for the paper "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos". 🏠 Project Page📜 arXiv 🧑‍💻 GitHub Sa2VA is the first unified model for the dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation… See the full description on the dataset page: https://huggingface.co/datasets/Dense-World/Sa2VA-Training.image-text-to-text8 likes611 downloads2mo agoHugging Face02bitersun /Sa2VA-finetune-exampleimagen<1K1 likes76 downloads11mo agoHugging Face03Dense-World /Sa2VA-Evalimage1 likes60 downloads1y agoHugging Face04tlzhang96 /Sa2VA-TrainingThis repository contains the code and data for the paper "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos". 🏠 Project Page📜 arXiv 🧑‍💻 GitHub Sa2VA is the first unified model for the dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation… See the full description on the dataset page: https://huggingface.co/datasets/tlzhang96/Sa2VA-Training.image-text-to-text0 likes13 downloads10mo agoHugging Face05albert6051 /llava_SA2VA_Refined_Garbageimagen<1K0 likes7 downloads2mo agoHugging Face06albert6051 /SA2VA_Refined_Garbageimagen<1K0 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.