VILA
Datasets
All datasets matching “VILA”MODUS-15Modality
MODUS — 15-Modality Aligned Dataset
MODUS is a large-scale, pixel-aligned 15-modality dataset for any-to-any
multimodal training. Every sample aligns 15 modalities covering appearance,
geometry, structure, segmentation, detection, text, and learned features.
Paper: https://huggingface.co/papers/2607.25948
Code: https://github.com/EPFL-VILAB/Modus
Modalities
Group
Modalities
Appearance
rgb, caption
Geometry
depth, normal
Structure
canny, sam_edge… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality.HR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.A2A-Video-examplesVACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/ali-vilab/VACE-Benchmark.MultihopSpatial
[ECCV 2026] MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Models
Project Page |
Paper |
Model
Overview
MultihopSpatial is a benchmark designed to evaluate whether vision-language models (VLMs) demonstrate robustness in multi-hop compositional spatial reasoning. Unlike existing benchmarks that only assess single-step spatial relations, MultihopSpatial features queries with 1 to 3 reasoning hops paired with… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/MultihopSpatial.holisafe-bench
⚠️ CONTENT WARNING: This dataset contains potentially harmful and sensitive visual content including violence, hate speech, illegal activities, self-harm, sexual content, and other unsafe materials. Images are intended solely for safety research and evaluation purposes. Viewer discretion is strongly advised.
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model (CVPR'26 Findings)
🌐 Website | 📑 Paper
📋 HoliSafe-Bench Dataset… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/holisafe-bench.
