datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LongVideoBenchBenchCheck-LongVideo
BenchCheck-LongVideo: frame-budget ladder for three open models (task 13b run package)
Run package for an agent on a separate GPU machine. Goal: on each of 30 long-video benchmarks
(mean video duration >= 300 s), answer the same up to 300 (155 to 320 multiple-choice items per benchmark, 8481 in total) multiple-choice items with THREE models at
four frame budgets, 32 / 128 / 512 / 1024 frames, at the model's own default resolution, and send
the per-item outputs back. The analysis… See the full description on the dataset page: https://huggingface.co/datasets/GMLRVigil/BenchCheck-LongVideo.LongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.
