tue-mps/videomt-dinov2-base-ytvis2019
07
VidEoMT-B on YouTube-VIS 2019
This repository contains the Hugging Face Transformers conversion of the official VidEoMT checkpoint yt_2019_vit_base_58.2.pth from tue-mps/VidEoMT.
Model details
- Architecture: VidEoMT with a DINOv2 ViT-B/14 with 4 register tokens backbone
- Task: video instance segmentation
- Dataset: YouTube-VIS 2019
- Input resolution: 640 x 640
- Number of frames: 2
- Paper: Your ViT is Secretly Also a Video Segmentation Model
Reported metrics
The metrics above are the numbers reported by the authors in the official model zoo.
Usage
from transformers import AutoModelForUniversalSegmentation, AutoVideoProcessor
model_id = "tue-mps/videomt-dinov2-base-ytvis2019"
processor = AutoVideoProcessor.from_pretrained(model_id)
model = AutoModelForUniversalSegmentation.from_pretrained(model_id)Use processor.post_process_instance_segmentation, processor.post_process_panoptic_segmentation, or processor.post_process_semantic_segmentation depending on the target task.
