datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gs-videos-v2PE-Video
PE Video Dataset (PVD)
[📃 Tech Report]
[📂 Github]
The PE Video Dataset (PVD) is a large-scale collection of 1 million diverse videos, featuring 120,000+ expertly annotated clips. The dataset was introduced in our paper "Perception Encoder".
Overview
PE Video Dataset (PVD) comprises 1M high quality and diverse videos. Among them, 120K videos are accompanied by automated and human-verified annotations. and all videos are accompanied with video description and keywords.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PE-Video.gs-videos-v3LSVQ-videosThis is an unofficial copy of the videos in the LSVQ dataset (Ying et al, CVPR, 2021), the largest dataset available for Non-reference Video Quality Assessment (NR-VQA); this is to facilitate research studies on this dataset given that we have received several reports that the original links of the dataset is not available anymore.
See FAST-VQA (Wu et al, ECCV, 2022) or DOVER (Wu et al, ICCV, 2023) repo on its converted labels (i.e. quality scores for videos).
The file links to the labels in… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LSVQ-videos.Videos_FL_0216generated-videosopen_video_datalongvideo_eval_videos
Long-RL: Scaling RL to Long Sequences (Evaluation Dataset - for research only)
Data Distribution
We strategically construct a high-quality dataset with CoT annotations for long video reasoning, named LongVideo-Reason. Leveraging a powerful VLM (NVILA-8B) and a leading open-source reasoning LLM, we develop a dataset comprising 52K high-quality Question-Reasoning-Answer pairs for long videos. We use 18K high-quality samples for Long-CoT-SFT to initialize… See the full description on the dataset page: https://huggingface.co/datasets/LongVideo-Reason/longvideo_eval_videos.Videos_FL_0204Videos_FL_0225FIRM-Video
FIRM-Video-SFT-90K
This repository releases the 90K SFT data for FIRM-Video.
The dataset covers three key evaluation dimensions:
Instruction Following (IF): whether the generated video accurately follows the text prompt.
Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity.
World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/FIRM-Video.VideoTemp-o3 VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
Illustration of the agentic pipeline in VideoTemp-o3. Given a video QA pair, the model performs on-demand temporal grounding to locate the most relevant segment, then refines it iteratively. Finally, it produces a reliable answer grounded in the pertinent visual evidence.
Data Source
The question and answer pairs used for training VideoTemp-o3 are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Keye/VideoTemp-o3.videomindvideo-diffusion-perceptiontrain_raw_video
ShareGPTVideo Raw ActivityNet Videos for Train data
All dataset and models can be found at ShareGPTVideo.
Contents:
Due to our scene split, we provide our processed activityNet videos corresponding to test frames in
train video frames
the processing script is process_activitynet.py
Videos_FL_0221Physics-aware-videos中文 README
🧱 Physics-aware Video Dataset
This dataset is a high-quality real-world video dataset focused on physical phenomena, designed for learning and evaluating physical laws from videos. It primarily covers the following classic physical processes:
🧱 Rigid-body motion / 🌊 Fluid dynamics
💫 Collision and rebound
💥 Explosion and burst phenomena
🌫️ Smoke, dust, and particle scattering
🌍 Gravity and inertia (falling, rolling, acceleration, etc.)
The dataset is constructed… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/Physics-aware-videos.VideoSSR-30kVoT-video-latent-archivemead_hdtf_400_merge_video_audio_frames_onlyaislop-videosHand-action-videos中文 README
✋ Hand Action Video Dataset
✋ Hand Action Video Dataset
This dataset is a real-world video dataset focused on hand actions (Hand Action Video Dataset), with an emphasis on common operations in hand–object interaction scenarios. It contains a large number of videos featuring:
✋ Clean and clear hand actions
🧱 Explicit hand–object interaction relationships
The dataset is constructed via video retrieval on **Datacube ** combined with automatic filtering using multimodal… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/Hand-action-videos.video-cVideoVista-CoTs
VideoVista-CoTs
This repository contains VideoVista-CoTs, used in Uni-MoE-2.0 training.
This dataset samples a portion of data from LLaVA-Video-178K, SEED-Bench-R1, SR-91K, and STAR, and uses our automatic Video QA generation framework to perform multi-step reasoning annotations for filtered complex questions.
The automatic video QA generation codes and our VideoVista series are presented in VideoVista Family
Citation
If you find VideoVista-CulturalLingo useful for your… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/VideoVista-CoTs.nsfw-video-still-caption-grid-onlyVideos_FL_0205galaxy_video_clip_featuresvideo_chatgpt_activitynet_videosqvhighlight_internvideo2_videoclip_6b_w2stmp-transfer-videoclips
