atharv6f/flash-attention-explorer
0
FlashAttention Explorer
Interactive demonstrations for the Attention Optimizations article series.
Features
- Visualizer: FlashAttention tiling animation and online softmax state tracking
- Benchmark: Compare Math, FlashAttention, and Memory-Efficient backends on real models
- Prefill vs Decode: Understand why FlashAttention helps prefill more than decode
- GQA/MQA: Explore Grouped-Query Attention and KV cache savings
- Memory Budget: Plan deployments with memory calculations
Models Used
Related Articles
This Space accompanies the "Attention Optimizations" series on HuggingFace:
- Standard Attention — The IO Problem
- FlashAttention — The Tiling Strategy
- FlashAttention — Online Softmax
- FlashAttention — IO Analysis
Technical Details
- Uses PyTorch SDPA (Scaled Dot Product Attention) with backend switching
- Zero GPU for real benchmarks
- Gradio Blocks for interactive UI
