datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LongVT-Source
LongVT-Source
This repository contains the source video and image files for the LongVT project.
Overview
LongVT is an end-to-end agentic framework that enables "Thinking with Long Videos" via interleaved Multimodal Chain-of-Tool-Thought. This dataset provides the raw media files referenced by the training annotations in LongVT-Parquet.
Dataset Structure
The source files are organized by dataset type and stored as zip archives:
Training Data… See the full description on the dataset page: https://huggingface.co/datasets/longvideotool/LongVT-Source.LongVideoBench
Dataset Card for LongVideoBench
Large multimodal models (LMMs) are handling increasingly longer and more complex inputs. However, few public benchmarks are available to assess these advancements. To address this, we introduce LongVideoBench, a question-answering benchmark with video-language interleaved inputs up to an hour long. It comprises 3,763 web-collected videos with subtitles across diverse themes, designed to evaluate LMMs on long-term multimodal understanding.
The… See the full description on the dataset page: https://huggingface.co/datasets/lccshunli/LongVideoBench.LongVideoBenchVideoSIAH-Eval
VideoSIAH-Eval
Evaluation benchmark for LongVT, containing 652 unique QA pairs across 244 long-form videos with human-in-the-loop validation.
Update (2026-03): The initial release contained 1,280 entries due to unintentional duplication during data export. This version has been cleaned to 652 unique QA pairs. Since each entry was an exact copy, all evaluation metrics reported in the paper remain unchanged.
LongVT-Parquet
LongVT-Parquet
This repository contains the training data annotations and evaluation benchmark for the LongVT project.
Overview
LongVT is an end-to-end agentic framework that enables "Thinking with Long Videos" via interleaved Multimodal Chain-of-Tool-Thought. This dataset provides the training annotations and evaluation benchmark in Parquet format, with source media files available in LongVT-Source.
Important Notes
For privacy reasons, media paths… See the full description on the dataset page: https://huggingface.co/datasets/longvideotool/LongVT-Parquet.LongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.LongVideoBenchlong_videolongvideoreflectionLongVideolongvideogen_t2v_memory_vimax_baselinelongvideogen_kling_v3_compact_5_vimax_wavespeed_kling_v3_5LongVT-Demo
