ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

Sparse Attention

2 posts tagged with "Sparse Attention"

August 12, 2026

MOD-DiT: Mixture of Distributions for Dynamic Sparse Attention in Video Diffusion Transformers

训练自由、无采样的动态稀疏注意力框架,通过线性近似模型预测三种核心注意力模式(块对角、平行对角、垂直)的强度演化,在 HunyuanVideo 上实现 2.05× 加速,在 Wan2.1 上实现 1.75× 加速。

Sparse Attention Video Diffusion Transformers Dynamic Masking Inference Acceleration Denoising Process Linear Approximation
August 12, 2026

[论文精读] OasisKV: Scaling In-Decode KV Cache

A memory-centric LLM inference system that decouples full KV-cache storage from HBM, using lookahead tokens from speculative decoding to prefetch only the most relevant KV blocks, achieving 1.69×-2.3× throughput gains within 0.7 points of full-attention accuracy.

KV Cache Sparsity LLM Inference Sparse Attention Speculative Decoding Prefetching Disaggregated Serving Memory Management

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.