hp-l33/sol-attn-experiments
0
Three slides: two same-prompt / seed video comparisons of fake64 and Full attention, followed by diffusion-loss curves for Full attention, sol-attn, PWT2, fake64, vsa style, and Shared slots.
The videos use complete 1344 × 768 × 124 outputs. Training updates the last two attention layers for 100 steps, with the new sparse runs using a cosine sparsity schedule from 80% to 95% over the first 50 updates. Sparse curves are evaluated at 95% sparsity; Full attention is dense. This is a small experiment with one training seed and two held-out videos, without repeated-run error bars.
Use the arrow keys or on-screen controls to navigate. Video playback is muted by default; each comparison has a synchronized playback button.
