mjuicem/StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.
1312k
Anomaly Context Understanding.zipdownload
Emotion Recognition.zipdownload
Misleading Context Understanding.zipdownload
Multimodal Alignment.zipdownload
Proactive Output_1-25.zipdownload
Proactive Output_26-50.zipdownload
Real-Time Visual Understanding_1-50.zipdownload
Real-Time Visual Understanding_101-150.zipdownload
Real-Time Visual Understanding_151-200.zipdownload
Real-Time Visual Understanding_201-250.zipdownload
Real-Time Visual Understanding_251-300.zipdownload
Real-Time Visual Understanding_301-350.zipdownload
Real-Time Visual Understanding_351-400.zipdownload
Real-Time Visual Understanding_401-450.zipdownload
Real-Time Visual Understanding_451-500.zipdownload
Real-Time Visual Understanding_51-100.zipdownload
Scene Understanding_1-25.zipdownload
Scene Understanding_26-50.zipdownload
Sequential Question Answering_1-25.zipdownload
Sequential Question Answering_26-50.zipdownload
Source Discrimination.zipdownload
