ddz16/CamInject-4B
022
CamInject-4B
Camera-movement understanding model that injects frozen VGGT camera tokens into Qwen/Qwen3-VL-4B-Instruct. Given a video, it outputs structured JSON describing every camera-movement segment.
- Paper: Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
- Project page: https://ddz16.github.io/cammotion.github.io
- Code: https://github.com/ddz16/CamDistill
⚠️ This model cannot be loaded with plain 🤗 Transformers. It requires a custom model type (registered via a plugin) and runs VGGT online to produce camera tokens. Loading it as a standard Qwen3VLForConditionalGeneration would not work correctly. Use the CamDistill repo.Usage
Clone the CamDistill repo and clone VGGT-Omega (set VGGT_OMEGA_REPO, see the repo's setup). CamInject runs VGGT online during inference:
VGGT_TEACHER_TYPE=vggt_omega \
python camera_movement_sft/infer_single.py \
--model ddz16/CamInject-4B \
--video /path/to/video.mp4 \
--variant caminjectSee the repo's README for environment setup and batch evaluation.
