ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

cuda-kernel

3 posts tagged with "cuda-kernel"

July 23, 2026

PRR: Predict, Reuse, and Repair — Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding

一种基于 EMA 预测器和在线 Softmax 增量修复的投机注意力运行时,打破 DSA 中 selection-to-attention 的关键路径依赖

llm-inference system-optimization sparse-attention cuda-kernel
July 23, 2026

PRR: Predict, Reuse, and Repair

Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding

long-context-llm inference-acceleration sparse-attention speculative-computation cuda-kernel flashattention kv-cache
July 23, 2026

PRR: Predict, Reuse, and Repair

Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding via Speculative Execution and Incremental Repair

long-context-llm sparse-attention kv-cache speculative-decoding cuda-kernel flashattention inference-acceleration dynamic-sparse-attention

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.