7 hours ago
- SANA-Video 2.0 is a hybrid video diffusion transformer at 5B and 14B scales, designed for high-quality 720p video generation on a single GPU, matching full-softmax transformers in quality with efficient linear attention scaling.
- To avoid quadratic attention, Hybrid Linear-Softmax Attention uses gated linear attention (O(N) complexity) with periodic gated-softmax anchors at a 3:1 ratio, restoring full-rank token interactions.
- Block Attention Residuals (AttnRes) route completed block summaries into later layers, boosting deep-layer effective rank by ~12% through anchor-feature reuse.
- From-scratch training learns the hybrid directly, with proxy studies establishing 25% softmax as optimal quality-efficiency trade-off.
- At 480p/40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s on a single H100, competitive with larger softmax models.
- Compiled DiT forward pass is 3.2× faster than full-softmax baseline at 720p/60s, with the gap enlarging for longer videos.
- Full-stack Sol-Engine optimization (kernel fusion, caching, sparse attention) further accelerates the backbone by 3.58×, making 720p/5s generation 120× faster than Wan 2.2-A14B on one H100.