August 5, 2026 ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression ResKV 将固定 KV 预算划分为精确主缓存和紧凑残差缓存,通过共享 softmax 重建被丢弃 token 的注意力贡献,显著提升长上下文推理效率 kv-cache kv-cache-compression long-context llm-inference attention cache-eviction decoding