shotbench
ShotBench
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
This is the official test set of ShotBench, comprising 3,572 question-answer pairs. Each sample is paired with either an image or a video clip. In total, ShotBench includes 3,049 images and 464 videos, primarily sourced from films that received Oscar nominations for Best Cinematography, ensuring high visual quality and strong cinematic style.
Paper: ShotBench: Expert-Level Cinematic Understanding in… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/ShotBench.shotbench-organized
ShotBench Dataset
Overview
ShotBench is a comprehensive benchmark for evaluating Vision-Language Models' understanding of cinematic language. It comprises 3,572 expert-annotated QA pairs from over 200 Oscar-nominated films.
Structure
— Complete dataset in JSON format
— Category distribution summary
— Per-category QA pairs
— 3,049 movie still images
— 464 movie video clips
8 Cinematography Dimensions
Dimension
Count
%… See the full description on the dataset page: https://huggingface.co/datasets/Baggio1012/shotbench-organized.ShotBench
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
This is the official test set of ShotBench, comprising 3,572 question-answer pairs. Each sample is paired with either an image or a video clip. In total, ShotBench includes 3,049 images and 464 videos, primarily sourced from films that received Oscar nominations for Best Cinematography, ensuring high visual quality and strong cinematic style.
Paper: ShotBench: Expert-Level Cinematic Understanding in… See the full description on the dataset page: https://huggingface.co/datasets/marvex/ShotBench.
