video-language
Video-LLaVAVideo-Benchmultimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.nepali-sign-language-video
Nepali Sign Language Video Dataset (NSL23)
First publicly available dataset for Nepali Sign Language (NSL) alphabet recognition.
Dataset Summary
Property
Value
Total videos
630
Total gestures
1,205
Consonants
36
Vowels
13
Volunteers
14 (5 experts, 9 beginners)
Environments
Bright, Dark, Prepared, Unprepared, RealWorld
Format
.mov video files
Splits
Split
Records
train
508
validation
61
test
61… See the full description on the dataset page: https://huggingface.co/datasets/Titung/nepali-sign-language-video.Sign_Language_Video_Key_points_DataCreative-Design-Studio-Interaction-Body-Language-Recognition-Video-Dataset
Creative Design Studio Interaction Body Language Recognition Video Dataset
In today's creative design industry, understanding the complex non-verbal communication among team members is crucial. However, existing body language recognition methods perform limitedly in dense interactive environments, struggling to accurately interpret subtle body movements and postures. This video dataset aims to tackle the technical challenges of body language analysis in creative discussions… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Creative-Design-Studio-Interaction-Body-Language-Recognition-Video-Dataset.
