vision-language-model
visionlanguagemodelogmultimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.ShareRobotSpatial-Blind-Spots-in-Vision-Language-Modelslicense: mit
model_evaluated:
name: Qwen3-VL-2B-Instruct
url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct
evaluation_notebook:
https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b
evaluation_setup: |
The model evaluated in this study is Qwen3-VL-2B-Instruct.
Evaluation was conducted using the Hugging Face Transformers library
with automatic device mapping (device_map="auto") and "bfloat16" dtype selection.
For each example:
The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.visionlanguagemodel
