fastvideo
Wan-Syn_77x448x832_600kWan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k.Mixkit-SrcHD-Mixkit-Finetune-HunyuanWan-Syn_77x768x1280_250kwan_t2v_distillation_dataset
