NJU-LINK/OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs ✨ Overview Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction. 🎬 OmniVideoBench fills this gap.It’s a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.
52.7k
No card is published for this repository, or it could not be fetched from Hugging Face right now.
