LongVideo
LongVideoBench
Dataset Card for LongVideoBench
Large multimodal models (LMMs) are handling increasingly longer and more complex inputs. However, few public benchmarks are available to assess these advancements. To address this, we introduce LongVideoBench, a question-answering benchmark with video-language interleaved inputs up to an hour long. It comprises 3,763 web-collected videos with subtitles across diverse themes, designed to evaluate LMMs on long-term multimodal understanding.
The… See the full description on the dataset page: https://huggingface.co/datasets/longvideobench/LongVideoBench.LongVideoDB-373K-VideosLongVideoBench
Dataset Card for LongVideoBench
Large multimodal models (LMMs) are handling increasingly longer and more complex inputs. However, few public benchmarks are available to assess these advancements. To address this, we introduce LongVideoBench, a question-answering benchmark with video-language interleaved inputs up to an hour long. It comprises 3,763 web-collected videos with subtitles across diverse themes, designed to evaluate LMMs on long-term multimodal understanding.
The… See the full description on the dataset page: https://huggingface.co/datasets/lccshunli/LongVideoBench.LongVT-Source
LongVT-Source
This repository contains the source video and image files for the LongVT project.
Overview
LongVT is an end-to-end agentic framework that enables "Thinking with Long Videos" via interleaved Multimodal Chain-of-Tool-Thought. This dataset provides the raw media files referenced by the training annotations in LongVT-Parquet.
Dataset Structure
The source files are organized by dataset type and stored as zip archives:
Training Data… See the full description on the dataset page: https://huggingface.co/datasets/longvideotool/LongVT-Source.LongVideoBenchLong-video-test-data
