CacheTune: Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
Frequency-guided and hardware-aware KV Cache reuse system achieving 3.72x-4.86x TTFT speedup and 3.93x-6.21x higher throughput via selective recomputation of semantic-critical tokens