VisionXLab/FIRM-Video
FIRM-Video-SFT-90K This repository releases the 90K SFT data for FIRM-Video. The dataset covers three key evaluation dimensions: Instruction Following (IF): whether the generated video accurately follows the text prompt. Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity. World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/FIRM-Video.
0428
1version https://git-lfs.github.com/spec/v12oid sha256:69f7a8e60021b1ecb1203c6882f673503be5e2090b103828df0a1b88b27100463size 107374182404 