ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

sequence-compression

1 post tagged with "sequence-compression"

July 18, 2026

Kwai Summary Attention (KSA):一种序列级 KV Cache 压缩的高效长上下文注意力机制

快手 OneRec 提出的序列级语义压缩注意力,通过可学习 summary token 实现 O(n/k) 的 KV Cache,长上下文精度匹配或超越 Full attention

llm-inference long-context efficient-attention kv-cache sequence-compression hybrid-attention

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.