rschaefer3333/fasth3-4step-preview
0
FastVideo FastH3 · 4-Step Preview — Showcase
An interactive showcase for FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree: a 35B text-to-audio-video model that produces synchronized video and audio from a text prompt in four transformer forwards, distilled with data-free DMD2 and VSA-H3 attention at 90% sparsity.
Why a showcase (not live inference)
The checkpoint is a 35B model whose tested inference path needs 4× NVIDIA B200 GPUs, FastVideo's custom VSA-H3 CUDA kernels, and CUDA 13. That is far beyond any hosted Space hardware (including ZeroGPU), so this Space explains the model, its 4-step architecture, and the exact commands to run it on your own Blackwell hardware, and links to the official blog for sample outputs.
Built with the huggingface-spaces skill.
