July 23, 2026 PRR: Predict, Reuse, and Repair Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding long-context-llm inference-acceleration sparse-attention speculative-computation cuda-kernel flashattention kv-cache
July 23, 2026 PRR: Predict, Reuse, and Repair Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding via Speculative Execution and Incremental Repair long-context-llm sparse-attention kv-cache speculative-decoding cuda-kernel flashattention inference-acceleration dynamic-sparse-attention