August 5, 2026 ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression ResKV 将固定 KV 预算划分为精确主缓存和紧凑残差缓存,通过共享 softmax 重建被丢弃 token 的注意力贡献,显著提升长上下文推理效率 kv-cache kv-cache-compression long-context llm-inference attention cache-eviction decoding
June 17, 2026 Flex Attention: A Programming Model for Generating Optimized Attention Kernels 编译器驱动的注意力编程模型,用几行PyTorch代码实现优化的注意力内核 attention compiler kernel system-optimization inference-serving torch-compile