datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FIRM-Video
FIRM-Video-SFT-90K
This repository releases the 90K SFT data for FIRM-Video.
The dataset covers three key evaluation dimensions:
Instruction Following (IF): whether the generated video accurately follows the text prompt.
Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity.
World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/FIRM-Video.video-diffusion-perceptionVideoSSR-30kVoT-video-latent-archivemead_hdtf_400_merge_video_audio_frames_onlynsfw-video-still-caption-grid-onlyVIDEOGENdecaf-rvos-davis17decaf-rvos-reasonvosvideo-diffusion-perception-feasibilitynsfw-video-still-caption-testultra-video-maskUnified-VideoDA-Generated-FlowsOptical flows associated with our work "We're Not Using Videos Effectively: An Updated Domain Adaptive Video Segmentation Baseline"
See the github for full instructions, but to install run
git lfs install
git clone https://huggingface.co/datasets/hoffman-lab/Unified-VideoDA-Generated-Flows
llava_video_max_256_frame_fps1rgba_videosnnnn_image_videosinsertion_dataset_expand_videoVideoBias
