Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
训练无关动态稀疏注意力:通过查询依赖高斯校准阈值、在线 Softmax 内联稀疏化和零阶泰勒代理分数复用,在不修改权重的情况下为视频生成 DiT 提供 2.0×–5.1× 端到端加速
6 posts tagged with "inference-acceleration"
训练无关动态稀疏注意力:通过查询依赖高斯校准阈值、在线 Softmax 内联稀疏化和零阶泰勒代理分数复用,在不修改权重的情况下为视频生成 DiT 提供 2.0×–5.1× 端到端加速
Training-free parallel decoding framework for block-wise diffusion LLMs using gated wavefront execution and heterogeneous wavefront packing
Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding
Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding via Speculative Execution and Incremental Repair
Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation
A speculative decoding runtime for stateful linear-attention models that verifies chains and trees with topology-aware kernels, stores compact factors to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter.