wan-2.2
Datasets
All datasets matching “wan-2.2”Wan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k.wan2.2_loraWan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k.wan2.2-Loraswan2.2llibero4in1_wan2.2vae_latent_dataset
LIBERO 4in1 Wan2.2-VAE Latent Cache
Pre-encoded latent tensors for LIBERO 4 suites under Cosmos Wan2.2-VAE.
Skip on-the-fly VAE encoding during training — load this cache directly.
Overview
Pre-encoded latent cache for LIBERO 4in1 benchmark (libero_spatial, libero_object, libero_goal, libero_10 — 4 suites × 10 tasks, ~1700 episodes total). Each raw video frame is encoded once with Wan2.2-VAE, then saved as .pt tensors for direct loading during action-policy… See the full description on the dataset page: https://huggingface.co/datasets/MangoGoes/libero4in1_wan2.2vae_latent_dataset.
