CoolFace
Datasetpublic

VisionXLab/FIRM-Video

FIRM-Video-SFT-90K This repository releases the 90K SFT data for FIRM-Video. The dataset covers three key evaluation dimensions: Instruction Following (IF): whether the generated video accurately follows the text prompt. Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity. World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/FIRM-Video.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes428downloads
Dataset Card

FIRM-Video-SFT-90K

This repository releases the 90K SFT data for FIRM-Video.

The dataset covers three key evaluation dimensions:

  • Instruction Following (IF): whether the generated video accurately follows the text prompt.
  • Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity.
  • World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and object/character structure.

For further details, please refer to the following resources:

  • Paper:
  • GitHub: https://github.com/VisionXLab/FIRM-Reward
  • Models&Benchmark: https://huggingface.co/collections/VisionXLab/firm-reward

Citation

bibtex